Appearance
PVE-1 ZFS Filesystem Migration & Re-imaging Runbook
Date: September 4, 2026
Cluster: clusterfuck (PVE 9.2.11)
Target Node: pve (PVE-1, 192.168.0.100)
1. Pre-Migration Node Audit Findings
| Component | Current State (PVE-1) | Target State Post-Reimage |
|---|---|---|
| Boot / OS Disk | 1TB SK hynix PC611 NVMe (/dev/nvme0n1) | 1TB SK hynix PC611 NVMe (/dev/nvme0n1) |
| OS Partition / FS | ext4 on LVM VG pve (/dev/pve/root) | ZFS on root (rpool/ROOT/pve-1) |
| Guest Storage | local-lvm (LVM-thin: pve/data, 762 GB) | local-zfs (rpool/data, matching PVE-2 & PVE-3) |
| ISO / Template Storage | local (dir, /var/lib/vz) | local (dir, /var/lib/vz on ZFS) |
| Mass Storage Array | 4x 4TB Seagate Enterprise (/dev/sda..sdd) in ZFS RAIDZ1 (tank) | Unchanged (tank re-imported cleanly) |
| Mass Storage Mount | /tank/media_root | /tank/media_root + /mnt/pve/tank_media symlink |
| Network Interfaces | eno1 (1G), enp2s0 (2.5G) | Same |
| Network Bonding | bond0 (active-backup, primary enp2s0) | Same |
| Bridges | vmbr0 (192.168.0.100/24, gw 192.168.0.1)vmbr1 (10.10.10.1/24) | Same |
| Active Guests | LXC 131 (docker-media), 132 (immich-docker), 200 (plex-jellyfin) | Temporarily migrated to PVE-2, then migrated back |
2. Container Migration to PVE-2 (LXC 132, 131, 200)
Why Direct Migration Previously Failed:
- Storage Type Mismatch: PVE-1 uses
local-lvm, whereas PVE-2 useslocal-zfs. Migrations must pass--target-storage local-zfs. - Local Bind Mount Check: LXC 131, 132, and 200 contain
mp0: /tank/media_root,mp=/mnt/media_root,acl=1. Becauseshared=1is missing, Proxmox aborts withcan't migrate local bind mount 'mp0'.
Preparation on PVE-2 (Ensure Media Mount Path Exists):
On PVE-2 (via SSH or Web Shell):
bash
# Ensure /tank/media_root symlink exists on PVE-2 pointing to the NFS mount:
mkdir -p /tank
ln -sfn /mnt/pve/tank_media /tank/media_rootExecution on PVE-1:
On PVE-1 (via SSH or Web Shell):
bash
# 1. Update mp0 on all three containers to declare the bind mount as shared:
pct set 132 -mp0 /tank/media_root,mp=/mnt/media_root,acl=1,shared=1
pct set 131 -mp0 /tank/media_root,mp=/mnt/media_root,acl=1,shared=1
pct set 200 -mp0 /tank/media_root,mp=/mnt/media_root,acl=1,shared=1
# 2. Stop LXC 131 (132 and 200 are already stopped):
pct stop 131
# 3. Migrate containers to PVE-2 converting storage to local-zfs:
pct migrate 132 pve2 --target-storage local-zfs
pct migrate 131 pve2 --target-storage local-zfs
pct migrate 200 pve2 --target-storage local-zfs
# 4. (Optional) Start them on PVE-2 if you want services online during the re-image:
pct start 132
pct start 131
pct start 2003. Host Configuration Backup Script (Run on PVE-1)
Run this command block directly on PVE-1 to save all configuration files into /root/pve1-config-backup/ and archive them to /root/pve1-config-backup.tar.gz:
bash
mkdir -p /root/pve1-config-backup
# 1. Network & Hosts
cp -a /etc/network/interfaces /root/pve1-config-backup/
cp -a /etc/hosts /root/pve1-config-backup/
cp -a /etc/resolv.conf /root/pve1-config-backup/
# 2. Storage & NFS Exports
cp -a /etc/pve/storage.cfg /root/pve1-config-backup/ 2>/dev/null || true
cp -a /etc/exports /root/pve1-config-backup/ 2>/dev/null || true
cp -a /etc/exports.d /root/pve1-config-backup/ 2>/dev/null || true
# 3. Cluster & Corosync
cp -a /etc/pve/corosync.conf /root/pve1-config-backup/ 2>/dev/null || true
# 4. User namespace mappings (critical for unprivileged LXCs)
cp -a /etc/subuid /root/pve1-config-backup/
cp -a /etc/subgid /root/pve1-config-backup/
# 5. Kernel modules & udev rules (QSV GPU pass-through)
cp -a /etc/modules /root/pve1-config-backup/ 2>/dev/null || true
cp -a /etc/modprobe.d/ /root/pve1-config-backup/ 2>/dev/null || true
cp -a /etc/udev/rules.d/ /root/pve1-config-backup/ 2>/dev/null || true
# 6. Dump active system states to text file
cat << 'EOF' > /root/pve1-config-backup/system_state.txt
=== PVE VERSION ===
$(pveversion -v)
=== ZFS POOLS ===
$(zpool status -v)
$(zfs list)
=== NETWORK STATE ===
$(ip addr)
$(ip route)
=== STORAGE CONFIG ===
$(cat /etc/pve/storage.cfg)
=== NFS EXPORTS ===
$(exportfs -v)
=== DISK LAYOUT ===
$(lsblk -o NAME,FSTYPE,SIZE,MOUNTPOINT,UUID,MODEL)
$(pvs)
$(vgs)
$(lvs)
EOF
# 7. Create archive
tar -czvf /root/pve1-config-backup.tar.gz -C /root pve1-config-backup
# 8. Copy to PVE-2 for safe keeping
scp /root/pve1-config-backup.tar.gz root@192.168.0.101:/root/4. Re-imaging PVE-1 Step-by-Step
Phase 1: Cluster Evacuation
- Confirm that LXC 131, 132, and 200 are on PVE-2 and all VMs/CTs on PVE-1 are cleared.
- Shut down PVE-1:bash
poweroff - On PVE-2 (verify cluster quorum remains 2/3):bash
pvecm status # Remove PVE-1 from Corosync: pvecm delnode pve
Phase 2: PHYSICAL SAFETY (Protecting tank)
CAUTION
MANDATORY SAFETY STEP: Before inserting the installer USB, physically disconnect the SATA/SAS data cables from the 4 Seagate 4TB HDDs (/dev/sda, /dev/sdb, /dev/sdc, /dev/sdd). Only the 1TB NVMe drive (/dev/nvme0n1) should be connected during the installation.
Phase 3: Proxmox VE Installation
- Boot from Proxmox VE 9 installer USB.
- Select Install Proxmox VE (Graphical).
- At the Target Harddisk screen:
- Click Options.
- Filesystem: Select
zfs (RAID0). - Harddisk 0: Select
/dev/nvme0n1(SK hynix 1TB NVMe). ashift:12compress:on(lz4 default).
- Management Network Configuration:
- Management Interface: Select
enp2s0(oreno1). - Hostname:
pve.rayweb.org(short name:pve). - IP Address:
192.168.0.100 - Netmask:
255.255.255.0(/24) - Gateway:
192.168.0.1 - DNS Server:
192.168.0.1
- Management Interface: Select
- Finish installation and reboot into the new PVE-1.
5. Post-Reinstall Configuration & Cluster Rejoin
Step 1: Restore Network Bond & Bridges
On fresh PVE-1 (/etc/network/interfaces):
ini
auto lo
iface lo inet loopback
iface eno1 inet manual
iface enp2s0 inet manual
auto bond0
iface bond0 inet manual
bond-slaves eno1 enp2s0
bond-miimon 100
bond-mode active-backup
bond-primary enp2s0
# Logic: "Use enp2s0 (2.5G) first. If it dies, switch to eno1 (1G) instantly."
auto vmbr0
iface vmbr0 inet static
address 192.168.0.100/24
gateway 192.168.0.1
bridge-ports bond0
bridge-stp off
bridge-fd 0
auto vmbr1
iface vmbr1 inet static
address 10.10.10.1/24
bridge-ports none
bridge-stp off
bridge-fd 0Apply network changes:
bash
ifreload -a || systemctl restart networkingStep 2: Re-join the Cluster
On PVE-1:
bash
pvecm add 192.168.0.101Verify cluster status:
bash
pvecm statusStep 3: Reconnect HDDs & Import tank
- Shut down PVE-1:
poweroff. - Reconnect the 4 SATA/SAS data cables to the 4 Seagate HDDs.
- Power on PVE-1.
- Import the pool:bash
zpool import tank zpool status tank - Re-add
tankstorage in Proxmox if needed (or verify viapvesm status).
Step 4: Re-enable NFS Export & Mount Path
On PVE-1:
bash
apt-get update && apt-get install -y nfs-kernel-server
# Restore /etc/exports:
echo "/tank/media_root 192.168.0.0/24(rw,sync,no_subtree_check,no_root_squash)" >> /etc/exports
exportfs -ra
systemctl restart nfs-kernel-server
# Bind mount /tank/media_root to /mnt/pve/tank_media (real directory, not symlink):
mkdir -p /mnt/pve/tank_media
echo "/tank/media_root /mnt/pve/tank_media none bind,x-systemd.requires-mounts-for=/tank/media_root 0 0" >> /etc/fstab
mount -a || mount --bind /tank/media_root /mnt/pve/tank_media
# Ensure POSIX ACLs are configured on ZFS dataset for unprivileged LXCs:
zfs set acltype=posixacl xattr=sa aclmode=passthrough tank/media_root
setfacl -R -m u:100000:rwx,d:u:100000:rwx,g:110000:rwx,d:g:110000:rwx /tank/media_rootStep 5: Migrate Containers Back to PVE-1
With PVE-1 now running native ZFS (local-zfs), migrate the containers back: On PVE-2:
bash
pct stop 132 131 200
pct migrate 132 pve
pct migrate 131 pve
pct migrate 200 pve
pct start 132
pct start 131
pct start 200Because all nodes now share the identical local-zfs storage identifier, migrations run natively via ZFS snapshot send/receive without manual storage remapping!