We All Have a Monitoring Checklist. Nobody Actually Follows It.
Quick story. A team I worked with had a 40-item production readiness checklist. Beautifully formatted. Color-coded priority levels. Lived in Notion. Got filled out during every launch review.
Their payment webhook processor stopped running on a Tuesday. They found out Friday afternoon — from the finance team.
Writing a checklist feels productive. You had an outage, you do a postmortem, you add "monitor X" to the list. Good. Except the list grows and nobody re-reads the whole thing. Items 1-5 get done because they're obvious. Items 15-40 are aspirational at best.
And the stuff that actually breaks in production? It's never on the list. It's the thing you didn't think of.
What Monitoring Should Actually Cover
Most setups nail the visible layer. HTTP endpoints, response times, error rates for your main API. That's table stakes. If you don't have that, stop reading and go set it up.
The gap is always in the invisible layer.
Scheduled tasks. Queue workers. Batch processes. The stuff that runs in the background, has no UI, and doesn't respond to HTTP requests. This is where outages hide for days. I've personally missed a failing cron job for over a week because nothing in our monitoring stack was watching it. The job would start, hit a config error, exit with code 0 (yeah, fun), and everyone assumed it was fine.
Instead of "what should we monitor," ask "what would hurt if it stopped working and we didn't know for 48 hours?"
Run that exercise with your team. You'll probably get a list like:
Email sending (transactional, not marketing)
Scheduled data imports/exports
The one Lambda function that processes invoices
Now check how many of those have active monitoring with clear alerting. In my experience? Maybe one or two.
The Dead Man's Switch Pattern
For anything that runs on a schedule, this pattern is probably the most reliable approach I've found.
The job runs. When it finishes successfully, it pings an external endpoint. If the ping doesn't arrive within the expected time window, you get an alert. Simple. No log parsing, no checking dashboards, no "did the cron even fire?"
You can build this yourself with a webhook and a timer. Or use an external service — there are several. Point is, the monitoring lives outside your infrastructure. If your entire server goes down, the alert still fires because the expected ping never arrives.
Alerts Need Owners, Not Channels
"Send all alerts to #ops-alerts" is how you get 500 unread messages and zero action.
Every alert should go to a person. A specific, on-rotation person who is responsible for acknowledging and acting on it. If nobody owns the alert, delete it. Seriously. An unowned alert is worse than no alert because it creates a false sense of coverage.
Btw, if your alert doesn't have a runbook attached — even a two-sentence one — it's just a notification that something is wrong with no guidance on what to do about it. That's noise, not signal.
Monitoring Doesn't Age Well
This one cost me a few hours at a previous job. We had monitoring set up for a cron job that ran every 6 hours. Someone changed it to hourly. The monitoring still expected a ping every 6 hours. So for 5 hours between pings, we had zero coverage. The job failed for 4 hours straight and we had no idea.
Config drift is real. Your monitoring setup is accurate on the day you configure it. After that, it's a snapshot of history. Unless you treat it like code — version it, review changes, update it alongside the systems it watches — it'll quietly become useless.
No checklist survives contact with a system that changes weekly. The habit of updating monitoring alongside every infrastructure change is worth more than any 40-item document in Notion.