Checking disk health
Disks usually warn you before they die. Nobody is listening, which is why the warning does not help.
Look at it
smartctl -a /dev/sda on Linux, or your controller's tool for disks
behind a RAID card. Run it once now so you know what healthy looks like on this
machine.
The values worth watching
- Reallocated sectors — zero is normal. Anything that keeps climbing is a disk on its way out.
- Pending and uncorrectable sectors — more serious than reallocated ones. These are reads that failed.
- Media and data integrity errors (NVMe) — same meaning.
- Percentage used / wear levelling (SSD and NVMe) — tells you how much write endurance is gone.
- CRC errors — usually a cable or a backplane rather than the disk itself. Worth a ticket too.
Temperature matters less than people think, but a disk that is suddenly ten degrees warmer than its neighbours is telling you something about airflow.
Run a self test
smartctl -t long /dev/sda and read the result a few hours later. A
short test misses what a long test finds.
When to open a ticket
As soon as the numbers move, not when the filesystem goes read-only. Send us the
full smartctl -a output and we schedule a replacement with you —
planned, at a time that suits you, instead of at three in the morning. See
when hardware fails.
Still stuck? Mail support@novogara.com — an engineer answers, at any hour. Back to the knowledge base