Appearance
🚑 Disaster Recovery & Drive Replacement Runbook
This runbook defines the emergency replacement and resilvering procedure for failed hard drives in the primary mass storage pool tank (ZFS RAIDZ1) on PVE-1.
🛡️ Drive Health Context & Standby Policy
- Known Historical Error: Drive
/dev/sda(Serial:ZC1B91PC, WWN:0x5000c500b5ede40f) has a recorded past read error (READ: 6in zpool status, 9 reallocated sectors, 4 uncorrectable). The error count is stable and not incrementing. - Run-Until-Death Policy: Drives in the array are intentionally kept running until they experience total physical failure. Do not trigger alerts or prompt for replacement during routine health checks.
- Cold Spare on Standby: A spare 4TB HDD is stored on-site and ready to replace any drive in the array immediately upon failure.
🛠️ Step-by-Step Drive Replacement Procedure
1. Identify the Failed Drive
Confirm the serial number, partition, and by-id symlink:
bash
# On PVE-1 host:
zpool status tank
ls -l /dev/disk/by-id/ | grep -E "ZC1B91PC|ZC1B9SWK|ZC1B9Q83|ZC1B938P"2. Offline the Failing Drive (Before Physical Removal)
Tell ZFS to take the device offline so the pool transitions to DEGRADED cleanly:
bash
# Example for ZC1B91PC:
zpool offline tank /dev/disk/by-id/ata-ST4000NM0035-1V4107_ZC1B91PC-part13. Physical Replacement
- Power down PVE-1 (or hot-swap if drive bays permit):bash
poweroff - Swap out the defective drive with the standby 4TB spare drive.
- Power on PVE-1.
4. Locate New Disk ID & Replace in Pool
- Find the
by-ididentifier for the newly inserted drive:bashls -l /dev/disk/by-id/ | grep ata- - Instruct ZFS to initiate replacement and begin resilvering:bash
zpool replace tank /dev/disk/by-id/<OLD_DISK_BY_ID> /dev/disk/by-id/<NEW_DISK_BY_ID> - Monitor resilvering progress:bash
zpool status tank 5
5. Clear Temporary Errors Once Resilvered
After resilvering finishes successfully (pool returns to ONLINE):
bash
zpool clear tank
zpool status tank