Skip to content

PVE-1 ZFS Filesystem Migration & Re-imaging Runbook ​

Date: September 4, 2026
Cluster: clusterfuck (PVE 9.2.11)
Target Node: pve (PVE-1, 192.168.0.100)


1. Pre-Migration Node Audit Findings ​

ComponentCurrent State (PVE-1)Target State Post-Reimage
Boot / OS Disk1TB SK hynix PC611 NVMe (/dev/nvme0n1)1TB SK hynix PC611 NVMe (/dev/nvme0n1)
OS Partition / FSext4 on LVM VG pve (/dev/pve/root)ZFS on root (rpool/ROOT/pve-1)
Guest Storagelocal-lvm (LVM-thin: pve/data, 762 GB)local-zfs (rpool/data, matching PVE-2 & PVE-3)
ISO / Template Storagelocal (dir, /var/lib/vz)local (dir, /var/lib/vz on ZFS)
Mass Storage Array4x 4TB Seagate Enterprise (/dev/sda..sdd) in ZFS RAIDZ1 (tank)Unchanged (tank re-imported cleanly)
Mass Storage Mount/tank/media_root/tank/media_root + /mnt/pve/tank_media symlink
Network Interfaceseno1 (1G), enp2s0 (2.5G)Same
Network Bondingbond0 (active-backup, primary enp2s0)Same
Bridgesvmbr0 (192.168.0.100/24, gw 192.168.0.1)
vmbr1 (10.10.10.1/24)
Same
Active GuestsLXC 131 (docker-media), 132 (immich-docker), 200 (plex-jellyfin)Temporarily migrated to PVE-2, then migrated back

2. Container Migration to PVE-2 (LXC 132, 131, 200) ​

Why Direct Migration Previously Failed: ​

  1. Storage Type Mismatch: PVE-1 uses local-lvm, whereas PVE-2 uses local-zfs. Migrations must pass --target-storage local-zfs.
  2. Local Bind Mount Check: LXC 131, 132, and 200 contain mp0: /tank/media_root,mp=/mnt/media_root,acl=1. Because shared=1 is missing, Proxmox aborts with can't migrate local bind mount 'mp0'.

Preparation on PVE-2 (Ensure Media Mount Path Exists): ​

On PVE-2 (via SSH or Web Shell):

bash
# Ensure /tank/media_root symlink exists on PVE-2 pointing to the NFS mount:
mkdir -p /tank
ln -sfn /mnt/pve/tank_media /tank/media_root

Execution on PVE-1: ​

On PVE-1 (via SSH or Web Shell):

bash
# 1. Update mp0 on all three containers to declare the bind mount as shared:
pct set 132 -mp0 /tank/media_root,mp=/mnt/media_root,acl=1,shared=1
pct set 131 -mp0 /tank/media_root,mp=/mnt/media_root,acl=1,shared=1
pct set 200 -mp0 /tank/media_root,mp=/mnt/media_root,acl=1,shared=1

# 2. Stop LXC 131 (132 and 200 are already stopped):
pct stop 131

# 3. Migrate containers to PVE-2 converting storage to local-zfs:
pct migrate 132 pve2 --target-storage local-zfs
pct migrate 131 pve2 --target-storage local-zfs
pct migrate 200 pve2 --target-storage local-zfs

# 4. (Optional) Start them on PVE-2 if you want services online during the re-image:
pct start 132
pct start 131
pct start 200

3. Host Configuration Backup Script (Run on PVE-1) ​

Run this command block directly on PVE-1 to save all configuration files into /root/pve1-config-backup/ and archive them to /root/pve1-config-backup.tar.gz:

bash
mkdir -p /root/pve1-config-backup

# 1. Network & Hosts
cp -a /etc/network/interfaces /root/pve1-config-backup/
cp -a /etc/hosts /root/pve1-config-backup/
cp -a /etc/resolv.conf /root/pve1-config-backup/

# 2. Storage & NFS Exports
cp -a /etc/pve/storage.cfg /root/pve1-config-backup/ 2>/dev/null || true
cp -a /etc/exports /root/pve1-config-backup/ 2>/dev/null || true
cp -a /etc/exports.d /root/pve1-config-backup/ 2>/dev/null || true

# 3. Cluster & Corosync
cp -a /etc/pve/corosync.conf /root/pve1-config-backup/ 2>/dev/null || true

# 4. User namespace mappings (critical for unprivileged LXCs)
cp -a /etc/subuid /root/pve1-config-backup/
cp -a /etc/subgid /root/pve1-config-backup/

# 5. Kernel modules & udev rules (QSV GPU pass-through)
cp -a /etc/modules /root/pve1-config-backup/ 2>/dev/null || true
cp -a /etc/modprobe.d/ /root/pve1-config-backup/ 2>/dev/null || true
cp -a /etc/udev/rules.d/ /root/pve1-config-backup/ 2>/dev/null || true

# 6. Dump active system states to text file
cat << 'EOF' > /root/pve1-config-backup/system_state.txt
=== PVE VERSION ===
$(pveversion -v)

=== ZFS POOLS ===
$(zpool status -v)
$(zfs list)

=== NETWORK STATE ===
$(ip addr)
$(ip route)

=== STORAGE CONFIG ===
$(cat /etc/pve/storage.cfg)

=== NFS EXPORTS ===
$(exportfs -v)

=== DISK LAYOUT ===
$(lsblk -o NAME,FSTYPE,SIZE,MOUNTPOINT,UUID,MODEL)
$(pvs)
$(vgs)
$(lvs)
EOF

# 7. Create archive
tar -czvf /root/pve1-config-backup.tar.gz -C /root pve1-config-backup

# 8. Copy to PVE-2 for safe keeping
scp /root/pve1-config-backup.tar.gz root@192.168.0.101:/root/

4. Re-imaging PVE-1 Step-by-Step ​

Phase 1: Cluster Evacuation ​

  1. Confirm that LXC 131, 132, and 200 are on PVE-2 and all VMs/CTs on PVE-1 are cleared.
  2. Shut down PVE-1:
    bash
    poweroff
  3. On PVE-2 (verify cluster quorum remains 2/3):
    bash
    pvecm status
    # Remove PVE-1 from Corosync:
    pvecm delnode pve

Phase 2: PHYSICAL SAFETY (Protecting tank) ​

CAUTION

MANDATORY SAFETY STEP: Before inserting the installer USB, physically disconnect the SATA/SAS data cables from the 4 Seagate 4TB HDDs (/dev/sda, /dev/sdb, /dev/sdc, /dev/sdd). Only the 1TB NVMe drive (/dev/nvme0n1) should be connected during the installation.

Phase 3: Proxmox VE Installation ​

  1. Boot from Proxmox VE 9 installer USB.
  2. Select Install Proxmox VE (Graphical).
  3. At the Target Harddisk screen:
    • Click Options.
    • Filesystem: Select zfs (RAID0).
    • Harddisk 0: Select /dev/nvme0n1 (SK hynix 1TB NVMe).
    • ashift: 12
    • compress: on (lz4 default).
  4. Management Network Configuration:
    • Management Interface: Select enp2s0 (or eno1).
    • Hostname: pve.rayweb.org (short name: pve).
    • IP Address: 192.168.0.100
    • Netmask: 255.255.255.0 (/24)
    • Gateway: 192.168.0.1
    • DNS Server: 192.168.0.1
  5. Finish installation and reboot into the new PVE-1.

5. Post-Reinstall Configuration & Cluster Rejoin ​

Step 1: Restore Network Bond & Bridges ​

On fresh PVE-1 (/etc/network/interfaces):

ini
auto lo
iface lo inet loopback

iface eno1 inet manual

iface enp2s0 inet manual

auto bond0
iface bond0 inet manual
	bond-slaves eno1 enp2s0
	bond-miimon 100
	bond-mode active-backup
	bond-primary enp2s0
# Logic: "Use enp2s0 (2.5G) first. If it dies, switch to eno1 (1G) instantly."

auto vmbr0
iface vmbr0 inet static
	address 192.168.0.100/24
	gateway 192.168.0.1
	bridge-ports bond0
	bridge-stp off
	bridge-fd 0

auto vmbr1
iface vmbr1 inet static
	address 10.10.10.1/24
	bridge-ports none
	bridge-stp off
	bridge-fd 0

Apply network changes:

bash
ifreload -a || systemctl restart networking

Step 2: Re-join the Cluster ​

On PVE-1:

bash
pvecm add 192.168.0.101

Verify cluster status:

bash
pvecm status

Step 3: Reconnect HDDs & Import tank ​

  1. Shut down PVE-1: poweroff.
  2. Reconnect the 4 SATA/SAS data cables to the 4 Seagate HDDs.
  3. Power on PVE-1.
  4. Import the pool:
    bash
    zpool import tank
    zpool status tank
  5. Re-add tank storage in Proxmox if needed (or verify via pvesm status).

Step 4: Re-enable NFS Export & Mount Path ​

On PVE-1:

bash
apt-get update && apt-get install -y nfs-kernel-server

# Restore /etc/exports:
echo "/tank/media_root 192.168.0.0/24(rw,sync,no_subtree_check,no_root_squash)" >> /etc/exports
exportfs -ra
systemctl restart nfs-kernel-server

# Bind mount /tank/media_root to /mnt/pve/tank_media (real directory, not symlink):
mkdir -p /mnt/pve/tank_media
echo "/tank/media_root /mnt/pve/tank_media none bind,x-systemd.requires-mounts-for=/tank/media_root 0 0" >> /etc/fstab
mount -a || mount --bind /tank/media_root /mnt/pve/tank_media

# Ensure POSIX ACLs are configured on ZFS dataset for unprivileged LXCs:
zfs set acltype=posixacl xattr=sa aclmode=passthrough tank/media_root
setfacl -R -m u:100000:rwx,d:u:100000:rwx,g:110000:rwx,d:g:110000:rwx /tank/media_root

Step 5: Migrate Containers Back to PVE-1 ​

With PVE-1 now running native ZFS (local-zfs), migrate the containers back: On PVE-2:

bash
pct stop 132 131 200
pct migrate 132 pve
pct migrate 131 pve
pct migrate 200 pve

pct start 132
pct start 131
pct start 200

Because all nodes now share the identical local-zfs storage identifier, migrations run natively via ZFS snapshot send/receive without manual storage remapping!


Authoritative operational repository and DR hub.