Skip to content

Automating server monitoring with cron + Slack

Skip the per-server monitoring subscription. One bash script and a cron job that post to a Slack webhook only when something actually breaks.

Every monitoring vendor wants $15 to $50 a month per server to send you a Slack message when a disk fills up. For one box that is fine. For eight, it is a line item you notice. We run a lot of small servers where a paid agent on each one makes no sense, so most of them report to Slack through a shell script and a cron entry. No agent, no dashboard, no per-node fee. Here is the exact setup we use, and the point where we stop and pay for a real tool instead.

The short version

Make a Slack incoming webhook. Write one bash script that checks disk, memory, load, and whether a couple of critical services are alive. Have it post to the webhook only when something crosses a threshold. Run it from cron every five minutes. That is the whole thing, and it covers maybe 80% of what a hobby-tier paid monitor would tell you. The other 20% (checks from outside your network, nice graphs, on-call rotation) is where a hosted tool earns its price, and we will get to that.

Why cron and a webhook, not an agent

A monitoring agent is a daemon that sits on the box, samples metrics, and ships them somewhere. That is genuinely useful when you have fleets of servers and want historical graphs. It is overkill when you have three VPSes and one question: “is anything on fire right now?”

The cron approach answers that one question with tools already on the server. No new daemon to keep alive, patch, or debug at 2am. If the script itself dies, cron just runs it again next tick. And because Slack’s incoming webhook is a single HTTPS POST, the whole alerting path is one curl command you can test by hand.

Step one: the Slack webhook

In Slack, create an app, turn on Incoming Webhooks, and add one to the channel you want alerts in. Slack hands you a URL that looks like https://hooks.slack.com/services/T00/B00/xxxx. Treat that URL like a password. Anyone who has it can post to your channel, so keep it in a file the script reads, not hard-coded and committed to a repo.

Test it before you build anything around it:

curl -X POST -H 'Content-type: application/json' \
  --data '{"text":"monitor test from web-01"}' \
  "$SLACK_WEBHOOK"

If the message lands in the channel, the hard part is done. Everything else is deciding what to check.

Step two: the check script

Here is a trimmed version of what we deploy. It checks disk on the root filesystem, available memory, one-minute load against core count, and whether nginx and mariadb are running. It alerts only on breach, so a healthy server stays silent.

#!/usr/bin/env bash
set -euo pipefail
SLACK_WEBHOOK="$(cat /etc/monitor/webhook)"
HOST="$(hostname -s)"

alert() {
  curl -sf -X POST -H 'Content-type: application/json' \
    --data "{\"text\":\":rotating_light: *$HOST* $1\"}" \
    "$SLACK_WEBHOOK" >/dev/null
}

# Disk: alert over 85% on /
disk=$(df --output=pcent / | tr -dc '0-9')
[ "$disk" -ge 85 ] && alert "disk at ${disk}% on /"

# Memory: alert under 200MB available
memfree=$(free -m | awk '/Mem:/ {print $7}')
[ "$memfree" -lt 200 ] && alert "only ${memfree}MB RAM available"

# Load: alert if 1-min load > 2x cores
cores=$(nproc)
load=$(awk '{print $1}' /proc/loadavg)
awk "BEGIN{exit !($load > $cores*2)}" && alert "load ${load} on ${cores} cores"

# Services
for svc in nginx mariadb; do
  systemctl is-active --quiet "$svc" || alert "$svc is DOWN"
done

Save it as /usr/local/bin/monitor.sh, chmod +x it, put the webhook URL in /etc/monitor/webhook with chmod 600, and adjust the service list per box.

Step three: cron, and the mistake everyone makes

Add it to root’s crontab:

*/5 * * * * /usr/local/bin/monitor.sh 2>&1 | logger -t monitor

Every five minutes, output piped to syslog so you have a trail. The mistake almost everyone makes here is alert flapping: the disk sits at 86%, and now you get a Slack ping every five minutes forever until you fix it. That noise trains you to ignore the channel, which defeats the entire point.

The fix is a simple state file. Only alert when a check crosses from healthy to broken, and send one “recovered” message when it crosses back. We keep a file per check under /run/monitor/ and compare the last state before posting. It adds maybe fifteen lines and turns a screaming channel into one that only speaks when something changed. If you build nothing else from this post, build that.

What this setup will not do

Be honest about the gaps, because they matter. First, it runs on the server it is watching. If the box goes fully offline (kernel panic, network drop, host outage) the script cannot tell you, because it is dead too. That is the big one. Second, there are no graphs and no history: you get “disk at 91% now”, not “disk has been climbing 2% a week for a month”. Third, there is no escalation. A Slack message at 3am pings a channel, but nobody’s phone rings.

The offline blind spot is fixable and worth fixing: add one external check. A free tier of an uptime service, or a second cheap VPS in another region running a curl against your site every minute, covers “is the whole box gone”. Internal script for the slow-burn stuff, one external ping for the hard-down case. That pairing catches the vast majority of real incidents.

When to stop scripting and pay for a tool

We do not pretend the DIY path scales forever. Move to a hosted monitor when you hit any of these:

  • More than eight to ten servers, where maintaining scripts across a growing fleet costs more time than the subscription saves.
  • You need graphs to argue with a hosting provider about their uptime, or to spot slow trends before they page you.
  • More than one person is on-call and you need real escalation: phone calls, rotations, acknowledgement.
  • Compliance or a client contract requires audited, retained monitoring data.

At that point the $15 a month is cheaper than your time, and a proper metrics stack pays off. Below that line, a shell script and a webhook are not a compromise. They are the right size for the job.

Where we come in

We set this up as part of onboarding on the servers we look after, tuned to each box’s real thresholds rather than generic defaults, plus the external check so a hard-down box never goes unnoticed. If you would rather not babysit crontabs and want someone watching the servers with you, that is what our ongoing support and monitoring covers. And if a box has already gone quiet and you are reading this after the fact, our writeup on how to investigate a server crash after it happens is the other half of the story. We fold SSL renewal into the same watch, so you never learn about a failed Let’s Encrypt renewal from an angry customer. Everything here runs from cron on purpose, because a stateless check does not mind overlapping with itself; if you are weighing up which scheduler a job should run on, that is the longer version of the argument.

Start here

Hit this one
yourself?

If any of the above is happening on your stack, send us the symptoms. We triage the same day and quote before we start.

Which layer is on fire?
Your details stay with us. Always.