Lesson 1/3
Monitoring & supervision: theory and good practice
Learn the fundamentals of IT monitoring: watching your infrastructure, the SNMP and WMI protocols, real-time dashboards, alert levels, agent vs agentless, and the mistakes to avoid. The full theory, with hard-won field experience, so you see failures coming.
☰
Contents
30
▾
0:03 Introduction to IT monitoring 0:29 What it is, and infrastructure sensors 0:52 A story: the hero vs preventive maintenance 1:16 Dashboards and seeing things in real time 2:03 Alerts, probes and sensors 2:13 Watching the CPU is not optional 2:30 RAM and swap, explained 3:10 Watching disk space, which is critical 3:29 Network monitoring and bandwidth 3:42 Types of monitoring: active vs passive 4:07 SNMP and why it works with everything 4:45 WMI and its limits 5:31 Agent vs agentless: the pros and cons 6:32 The popular tools: PRTG, Zabbix, Nagios 7:18 Alert levels: info, warning, critical 8:05 How the levels escalate, with a disk example 8:15 Alert channels: email, SMS, Teams 9:17 The danger of false positives 10:17 Calibrating your alert thresholds 11:02 Setting realistic thresholds from what you observe 12:57 Escalation and documenting your alerts 13:25 Web server metrics and certificates 14:02 Database metrics 14:25 Mail server metrics 14:41 Active Directory metrics 15:11 The mistake: monitoring far too much 15:48 Why documenting the process matters 16:21 Monitoring the monitoring tool itself 16:56 Predictive monitoring: anticipating rather than reacting 17:26 Conclusion and call to action