Independent reviewsWe host no filesNot affiliated with any vendor listed
Scannethub

Guide · covers Zabbix

Cutting alert noise: thresholds, dependencies and flapping

How to stop a monitoring system from paging for nothing: trigger hysteresis, dependencies, maintenance windows and flap detection in Zabbix, Icinga and Checkmk.

Picture an on-call phone receiving 214 notifications in four minutes at 03:12 on a Sunday. The cause is a single core switch reloading after a scheduled firmware push that someone forgot to put in the maintenance calendar. Every host behind it goes “unreachable”, every service on those hosts goes “critical”, and then everything recovers and notifies again. Nobody reads message 215. That is how alert fatigue starts: not with one bad threshold, but with a system that has no idea which failures are causes and which are consequences.

This guide covers the four levers that remove most of the noise in a self-hosted monitor: hysteresis on thresholds, dependencies between checks, maintenance windows, and flap detection. Examples use Zabbix syntax (7.0 LTS), with notes for Icinga, Nagios Core and Checkmk.

Step 1: measure the noise before changing anything

You cannot tune what you have not counted. Pull the last 30 days of notifications and group them:

  1. By trigger or service name — the top 10 usually account for more than half of all pages.
  2. By outcome — did anyone act? If a trigger fired 80 times and produced zero tickets, it is either wrong or should not notify.
  3. By duration — problems that resolve in under five minutes on their own are candidates for delay or hysteresis.

In Zabbix, Reports → Top 100 triggers gives the first list directly. In Icinga Web, filter the history view by notification type. Keep this list; you will compare against it after tuning.

Step 2: add hysteresis to thresholds

A threshold without hysteresis flips every time a metric wobbles around the line. CPU at 89%, 91%, 89%, 92% produces two problems and two recoveries in four minutes. The fix is to make the problem condition and the recovery condition different.

In Zabbix, use a recovery expression alongside the problem expression, and base both on a time window rather than one sample:

Problem expression:
  avg(/Linux by Zabbix agent/system.cpu.util,5m)>90

Recovery expression:
  avg(/Linux by Zabbix agent/system.cpu.util,10m)<75

The problem needs five minutes of sustained load; the recovery needs ten minutes below 75%. For “is it down” checks, use min() or max() over a window — max(/host/net.tcp.service[http],3m)=0 means three minutes of consecutive failures, not one lost probe.

Nagios and Icinga achieve the same effect with soft and hard states. Set max_check_attempts in a Nagios service template or an Icinga 2 template to 3–5 and use a shorter retry_interval: the service must fail several consecutive rechecks before it becomes a hard state and notifies.

template Service "generic-service" {
  max_check_attempts = 4
  check_interval = 1m
  retry_interval = 30s
}

Checkmk exposes this as the “Maximum number of check attempts for service” rule, plus separate WARN/CRIT levels on most checks.

Step 3: model dependencies so causes suppress consequences

The 214-message night was a dependency failure. If the monitor knows that 40 hosts sit behind one switch, it can report “switch down” and hold back the rest.

In Zabbix, open the host-unreachable trigger on each downstream host (or, better, on the template) and add the upstream device’s trigger under Dependencies. When the switch’s “Unavailable by ICMP ping” is in problem state, dependent triggers do not generate events. Do this with templates per site so new hosts inherit the chain.

In Nagios Core, the parents directive on host objects builds a reachability tree; when a parent is down, children become UNREACHABLE rather than DOWN, and you can skip notifying on u. Icinga 2 uses explicit Dependency objects with apply rules:

apply Dependency "uplink-switch" to Host {
  parent_host_name = "sw-core-01"
  disable_checks = false
  disable_notifications = true
  assign where host.vars.site == "hq" && host.name != "sw-core-01"
}

Checkmk builds parent relationships from its “Parents” host setting and can scan them automatically with its parent scan feature.

Do not stop at network topology. Services depend on each other too: if the database is down, the five application health checks behind it only add noise.

Step 4: use maintenance windows for everything planned

Planned work should never page anyone. The trick is making maintenance easy enough that people actually use it.

  • In Zabbix 7.0, maintenance lives under Data collection → Maintenance. Choose with data collection for most work so graphs keep a record; choose no data collection only when the data would be garbage. In the trigger action, enable Pause operations for suppressed problems so notifications wait until maintenance ends — and only fire if the problem is still there.
  • In Icinga and Nagios, schedule downtime on the host and its services (Icinga Web’s “schedule downtime” with “all services” ticked), or use ScheduledDowntime objects for recurring windows like Sunday patching.
  • Script it. A pre-deploy hook in your automation that calls the Zabbix API (maintenance.create) or the Icinga 2 REST API (/v1/actions/schedule-downtime) for the affected hosts removes the “forgot to set maintenance” failure mode completely.

Step 5: detect and damp flapping

Flapping is a state that alternates rapidly: a Wi-Fi bridge dropping every few minutes, a disk sitting on the threshold, a service restarting under a watchdog.

Nagios Core and Icinga 2 have built-in flap detection. Nagios tracks the percentage of state changes over the last 21 checks, Icinga 2 uses a similar rolling window of recent states, and both mark the object as flapping above a high threshold, clearing it below a low one. In Icinga 2, the defaults are flapping_threshold_high = 30 and flapping_threshold_low = 25; enable it per object with enable_flapping = true. While flapping, individual state-change notifications are held back and one “flapping started” message goes out instead.

Zabbix has no dedicated flap state, so model it with a trigger that counts changes:

changecount(/host/net.if.status[eth0],30m)>6

Give that trigger its own severity and route it to a ticket queue rather than a pager. Then make the underlying up/down trigger depend on it, so a flapping link produces one problem instead of dozens.

Step 6: fix routing, not just triggers

Even perfect triggers are noise if they reach the wrong people:

  1. Separate severities by channel. Disaster/High → pager. Average → team chat. Warning/Information → dashboard only.
  2. Add a delay to the first notification for non-critical severities. Zabbix action escalation steps can start at step 2 after 10 minutes; Nagios has first_notification_delay; Checkmk has a “Delay service notifications” rule.
  3. Deduplicate on the receiving side. If you forward to an incident tool, group by host and time window.

Common mistakes

  • Raising thresholds until the alert goes quiet. That hides real problems; use windows and hysteresis instead.
  • Dependencies defined per host by hand. They rot. Define them in templates or apply rules.
  • Maintenance without data collection by default. You lose the graphs that explain what happened.
  • Checking every 10 seconds “to be safe”. Faster polling multiplies flaps and database load; see how big your monitoring server should be.

Noise handling is one of the biggest real differences between platforms — compare them in Zabbix vs Checkmk and Icinga vs Nagios, or browse the server and service monitoring category for the wider field.