How to investigate a server crash post-mortem
A server that crashed once will almost certainly crash again unless you find out why. That is the whole point of a post-mortem: not to assign blame, but to turn one bad night into a written cause and a fix that holds. Most of the pages that rank for a server post-mortem investigation are forum threads or generic “top causes of server crashes” lists. Useful for a hunch, useless at 3 a.m. when a client’s store is down and you have to actually find the reason.
Here is the runbook we follow after every crash we get called into. It is log-first, it is ordered, and it assumes you have shell access and roughly an hour before someone asks for an answer.
First, define what actually happened
“The server crashed” covers four very different failures, and each one leaves evidence in a different place. Before you touch a log, decide which one you are chasing:
- A single service died (nginx, MySQL, PHP-FPM) but the box stayed up.
- The kernel panicked and the machine rebooted itself.
- Resources ran out (RAM, disk, or file descriptors) and the OOM killer started shooting processes.
- The host or hypervisor pulled the plug, and nothing on the guest is at fault at all.
Run uptime first. If the box has been up for three minutes and the incident was an hour ago, it rebooted, and you are looking at a panic, a hardware event, or a host action. If uptime spans the whole incident, the machine never went down and a service is your culprit. This one command saves you from reading the wrong logs for twenty minutes.
Read the logs in the right order
The timeline is your best friend. Line up timestamps from three sources and the story usually tells itself. Start with the kernel ring buffer and the system journal, narrowed to the incident window:
journalctl --since "2026-07-15 02:00" --until "2026-07-15 03:00"— everything the system recorded in that hour.journalctl -k -b -1— kernel messages from the previous boot, which survive a reboot and often hold the panic trace.dmesg -T | grep -i -E "oom|killed process|panic|error"— the loud events, with human-readable timestamps.
If you see “Out of memory: Killed process,” you have your answer and you can skip straight to the memory section. If the previous boot’s kernel log ends mid-sentence with no shutdown message, that is a hard lockup or a host-side kill, not a software crash on your side.
Rule out the boring causes before the exotic ones
The single most common “mystery crash” we get handed is a full disk. A database that cannot write its files will stop accepting connections, and the symptom looks like a database crash when the real problem is df showing 100%. A database that is slow rather than stopped is a different animal, usually a MariaDB config that never got tuned on CyberPanel. Check disk and inodes before anything clever:
df -handdf -i— space and inodes, because you can run out of inodes with gigabytes free.free -mplus the OOM check above — memory pressure.grep -i error /var/log/mysql/error.log(or the equivalent for your service) — the service’s own account of its death.
In our incident data over the last two years, roughly two out of three “the server crashed” calls come down to a full disk, a runaway process eating RAM, or a bad config pushed in a deploy an hour earlier. Hardware and kernel bugs are real, but they are the minority. Check the cheap causes first.
Tie the cause to a change
Servers rarely fail at random. Something moved: a deploy, a package update, a cron job, a traffic spike. Once you have a suspected cause, look for what changed near the timestamp. last reboot gives you the reboot history. Your package manager log (/var/log/dpkg.log or /var/log/yum.log) tells you if an update landed just before. Your deploy log or Git history tells you if code shipped. Correlation is not proof, but a crash that starts ninety seconds after a config reload is not a coincidence.
Write it down while it is fresh
The post-mortem is a short document, not a saga. We keep ours to five fields: what happened, the timeline, the root cause, the immediate fix, and the change that stops a repeat. That last field is the only one that matters in six months. “Restarted MySQL” is not a fix; “added disk alerting at 80% and moved backups off the data volume” is. If the write-up does not end with a change to configuration, monitoring, or process, the same crash is already scheduled.
We run this exact process as part of our Linux server restoration work, and we hand the written post-mortem to the client at the end. If you would rather not be the one reading kernel logs at 3 a.m., that is what we are for. If the box has been unstable for a while and you want to get ahead of it, the five signs your VPS needs an audit is a good gut-check, and the Linux server hardening checklist covers the preventative side once the immediate fire is out.
The short version
Catching the next crash sooner is far easier when a lightweight cron-and-Slack monitor is already watching disk, memory, and services for you.
Check uptime to learn whether the box rebooted. Read the kernel log and journal for the incident window. Rule out full disk and out-of-memory before anything exotic. Tie the cause to a change. Then write five fields and, most importantly, the one change that keeps it from happening twice. Do that every time and your servers get quieter, which is the only metric that counts.
Next in the journal
- 15 Jul 2026 ChatGPT for product descriptions: a workflow that scales ChatGPT can write a product description in four seconds. It can also write four hundred descriptions that all sound like the same mildly enthusiastic…
- 14 Jul 2026 Five WordPress backup methods I tested: real restore times A backup you have never restored is a guess, not a backup. Most guides stop at “set up automated backups and store them offsite,”…