“Two vCPUs and 4 GB of RAM should be plenty” is how most self-hosted monitoring servers begin, and for the first month it usually is. Then someone imports the full network-interface template for 60 switches, the history table passes 200 GB, the housekeeper starts running for six hours a night, and the web frontend takes twenty seconds to open a graph. The server was not undersized on day one. It was sized for a number nobody had actually calculated.
This guide shows how to calculate it. The worked numbers use Zabbix, because its data model makes the variables explicit, but the same reasoning applies to any database-backed monitor. Notes for RRD-based tools like LibreNMS and Checkmk are at the end.
The variables that matter
Forget host counts for a moment. A host with a ping check and a host with 400 SNMP interface counters are not the same load. What drives sizing:
- Items — every individual metric being collected (one interface’s inbound octets is one item).
- Update interval per item.
- NVPS (new values per second) — the sum of
1 / intervalacross all enabled items. This is the single number that best predicts CPU, write IOPS and database growth. Zabbix shows it on the dashboard as “Required server performance, new values per second”. - History retention — how many days raw values are kept.
- Trend retention — how long hourly aggregates (min/avg/max/count) are kept. Trends only exist for numeric items.
- Housekeeping method — row deletes versus dropping partitions or chunks.
- Topology — whether proxies collect data on the server’s behalf.
Step 1: count items and compute NVPS
Build a small table per template: number of hosts using it, items per host, and the typical interval. A spreadsheet is fine. Remember that low-level discovery multiplies items — a 48-port switch with an interface template collecting eight metrics per port is 384 items before you add fans, PSUs and CPU.
NVPS = Σ (items_in_group / interval_seconds)
Step 2: estimate history and trends storage
Zabbix’s own documentation uses rough per-row sizes for estimates. Real numbers depend on database engine, indexes and compression, so treat these as planning figures, not promises:
- History: about 90 bytes per stored value.
- Trends: about 90 bytes per item per hour.
- Events: roughly 250 bytes per event, usually minor next to history.
history_bytes/day = NVPS × 86,400 × 90
trends_bytes/year = numeric_items × 24 × 365 × 90
A worked example
Take a mid-sized hypothetical estate with two offices:
| Group | Hosts | Items/host | Interval | Items | NVPS |
|---|---|---|---|---|---|
| Linux servers (agent) | 120 | 150 | 60 s | 18,000 | 300 |
| Windows servers (agent) | 60 | 180 | 60 s | 10,800 | 180 |
| Access switches (SNMP) | 70 | 300 | 120 s | 21,000 | 175 |
| Core/firewalls (SNMP) | 8 | 500 | 60 s | 4,000 | 67 |
| Inventory-type items | all | 20 | 3,600 s | 5,160 | 1.4 |
| Total | 258 | 58,960 | ≈ 723 |
Round up to 750 NVPS for headroom.
- History, 14 days: 750 × 86,400 × 90 ≈ 5.8 GB per day, so about 82 GB for 14 days.
- Trends, 1 year: assume 55,000 numeric items. 55,000 × 24 × 365 × 90 ≈ 43 GB per year.
- Indexes, events, config, WAL/binlogs and free space for maintenance: add 50–100%.
So a database volume of roughly 250–300 GB on SSD or NVMe covers a year comfortably. Keeping history for 90 days instead of 14 would push history alone past 500 GB — which is why the usual advice is short history, long trends.
For compute, Zabbix’s published hardware guidance places an installation of this size in its “medium to large” band. As a rough rule of thumb, a server with 8 vCPUs and 32 GB RAM, with the database on the same host, has enough headroom for 750 NVPS; above roughly 1,500–2,000 NVPS the usual recommendation is to move the database to its own machine. Treat these as planning figures, not promises, and check the current requirements page for your Zabbix version before sizing hardware.
Step 3: tune the server processes and caches
Default zabbix_server.conf values suit a demo, not 750 NVPS. The settings to revisit:
CacheSize=256M # configuration cache: hosts, items, triggers
HistoryCacheSize=128M # buffer between collectors and DB writers
TrendCacheSize=64M
ValueCacheSize=512M # recent values used by trigger functions
StartDBSyncers=4
StartPollers=20
StartPollersUnreachable=4
StartPingers=4
Zabbix 7.0 added asynchronous pollers for agent, SNMP and HTTP checks (StartAgentPollers, StartSNMPPollers, StartHTTPAgentPollers), which handle many concurrent requests per process, so the classic “raise StartPollers until busy drops” routine matters less than it used to. Either way, the rule is the same: watch the internal items for process utilization and cache free space, and add capacity when collectors sit above ~75% busy.
Step 4: decide how old data gets removed
The built-in housekeeper deletes expired rows with DELETE statements. On a history table with hundreds of millions of rows, that is slow, generates heavy I/O, and on MySQL can leave fragmented tables.
Two better options:
- PostgreSQL with TimescaleDB. Zabbix supports it natively; history and trends become hypertables, expiry drops whole chunks, and compression of older chunks can cut storage substantially. This is our default for new builds.
- MySQL/MariaDB with partitioning. Partition history and trend tables by day or month using a maintained partitioning script, then disable the housekeeper for history and trends in Administration → Housekeeping. Dropping a partition takes seconds.
Whichever you pick, set it up before go-live. Migrating a 300 GB table to partitions later means a long maintenance window.
Step 5: add proxies before the server needs them
A Zabbix proxy collects data for a site or segment, buffers it locally, and sends it to the server in batches. Proxies:
- reduce the number of connections and pollers on the central server;
- keep collecting during WAN outages and backfill afterwards;
- let you reach networks the server cannot route to directly.
Use one per remote site, plus one per large segment in the main datacenter once you pass a few hundred NVPS. Active proxies (proxy connects to server) are easier through firewalls. Zabbix 7.0 also supports proxy groups for load balancing and failover.
RRD-based tools are different
LibreNMS and Checkmk store metrics in RRD files, which are fixed-size ring buffers: disk usage grows with the number of devices and ports, not with time. Sizing there is about write IOPS (thousands of small file updates every polling cycle), so use SSDs, run rrdcached, and add distributed pollers when a polling cycle approaches its interval. LibreNMS polls every five minutes by default; if a full poll run takes more than about four of those minutes, you are out of headroom.
Common mistakes
- Sizing by host count. Always compute NVPS.
- Keeping 90 days of raw history “just in case”. Trends answer almost every capacity question.
- Running the database on spinning disks or burstable cloud volumes. Write latency becomes a backlog in the history cache.
- Leaving the housekeeper on with partitioning enabled. It will try to delete rows you already dropped.
- Polling too often. Halving an interval doubles NVPS; see cutting alert noise for why it rarely improves detection.
Related reading
For the platform decision itself, see Zabbix vs Checkmk and the open-source monitoring category. Packages and repository signing keys should come from the vendor; our where to get it page covers verification.