When memory runs out
A process disappears without a trace and nobody knows why. Nine times out of ten the kernel killed it, and it wrote down exactly why.
Find out what happened
dmesg -T | grep -i -E 'killed process|out of memory' or
journalctl -k | grep -i oom. You get the process, the time and how much
memory it was holding. That is usually the whole investigation.
Why the kernel kills instead of slowing down
When memory is genuinely gone, Linux picks the process whose death frees the most and ends it. It does not warn, it does not ask. Your database being the biggest process on the machine makes it the most likely victim, which is why an OOM event so often looks like “the database crashed”.
Swap: yes, but modestly
A few gigabytes of swap gives the kernel room to move idle pages out and buys you
time to notice. What swap does not do is make up for missing memory: a server that
lives in swap is slower than one that restarts. Tune
vm.swappiness down to 10 for a database host, leave it alone otherwise.
Stop it happening again
- Cap the big consumers explicitly: the database buffer pool, the PHP or Java heap, worker counts. Defaults assume they own the machine.
- Count your workers. Thirty processes at 300 MB is nine gigabytes, and that is how a “sudden” shortage usually arrives.
- Run the services that matter under systemd with a restart policy, so an OOM kill is a blip instead of an outage.
- Alert on memory before it runs out, not after.
Or just add memory
Memory is the cheapest upgrade in a dedicated server and the one that most often fixes the actual problem. Ask us; on most builds it is a short maintenance window and a new price, not a migration.
Still stuck? Mail support@novogara.com — an engineer answers, at any hour. Back to the knowledge base