Skip to content

homelab-ops ​

Authoritative operational repository and documentation hub for the homelab infrastructure: 3-node Proxmox VE cluster (clusterfuck), ZFS storage architecture (tank), Docker service stacks, remote ingress security, automated backup pipelines, and disaster recovery procedures.


🖥️ Infrastructure at a Glance ​

The homelab runs a 3-node Proxmox VE cluster (PVE 9.2.11) backed by a dedicated Proxmox Backup Server (PBS):

NodeHostname / IPRole / Hardware SpecsStorage PoolsWorkload Strategy & Active Services
PVE-1pve
192.168.0.100
• Intel Core i7-10700 (8C/16T @ 2.90GHz)
• 32 GB DDR4 RAM (33.4 GB usable)
• Dual-NIC Active-Backup (bond0)
• local-zfs (~956 GB NVMe)
• tank (ZFS RAIDZ1, 16 TB raw)
• Cold Backup: 2TB WD HDD (/dev/sdf)
• local (dir)
Primary Compute & Storage Host (3-Domain Model):
• VM 199: homeassistant (HAOS, 1,757 entities)
• LXC 183: docker-edge (11 containers: Traefik, Cloudflare Tunnel, Auth, Glance, Termix, Dockhand Central, Tailscale)
• LXC 131: docker-media (31 containers: Arrs, Streaming, Immich, Hawser Agent, Backrest)
• LXC 129: docker-admin (16 containers: n8n, Mealie, Obsidian, Monitoring, Hawser Agent)
• LXC 133: bambuddy • LXC 180: ansible
• LXC 182: antigravity
• (Retired / Standby: LXC 185 edge-lxc, LXC 132 immich-docker, LXC 200 plex-jellyfin)
PVE-2pve2
192.168.0.101
• AMD Ryzen Embedded R2314 (4C/4T @ 1.40GHz)
• 16 GB DDR4 RAM (14.5 GB usable)
• Dual-NIC Active-Backup (bond0)
• local-zfs (242 GB NVMe)
• tank_media (NFS4 from PVE-1)
• local (dir)
HA Failover & Standby Node:
• Cluster quorum & HA manager (pve-ha-crm)
• Ready for compute failover
PVE-3pve3
192.168.0.102
• AMD Ryzen Embedded R2314 (4C/4T @ 1.40GHz)
• 16 GB DDR4 RAM (14.5 GB usable)
• Single NIC (nic0 1G) + WiFi
• local-zfs (242 GB NVMe)
• tank_media (NFS4 from PVE-1)
• local (dir)
HA Failover & Standby Node:
• Cluster quorum & golden templates (LXC 100)
• Ready for compute failover
PBSPBS
192.168.0.103
• Proxmox Backup Server
• Datastore: backup@pbs
• Automated weekly power cycle
• backup (PBS storage)Weekly Automated Backup Host:
• Powered on via Home Assistant Sun 00:45 AM
• Cluster backups run Sun 01:00 AM
• Auto-shutdown ~3h post-boot (task guard)
RouterFlint 2
192.168.0.1
• GL.iNet Flint 2 (GL-MT6000, OpenWrt)
• Automated WAN Watchdog & Self-Healing
• Local flash storageNetwork Gateway:
• internet_watchdog.sh runs every 5 min
• Auto-bounces WAN & sends Telegram alerts

⚡ Key Architectural Standards ​

  1. Universal Media Storage Mount: All media containers across all nodes use /mnt/pve/tank_media,mp=/mnt/media_root,shared=1 (without acl=1). On PVE-1, this is a local bind mount; on PVE-2/3, it is an NFS4 mount. See Universal Media Storage Runbook.
  2. LXC Permissions & UID/GID Mapping: Unprivileged LXCs map permissions via nas_shares group (GID 10000 in container $\rightarrow$ host GID 110000). Host ZFS dataset enforces POSIX ACLs granting full access to rayman-ssh and admin-ssh. See LXC Permissions Guide.
  3. Dual-Tier Remote Ingress & Security: External browser GUI access to Home Assistant and Proxmox is routed securely via Cloudflare Tunnels (cloudflared in Docker on docker-edge), Traefik, and TinyAuth/Pocket-ID OIDC (docker-edge LXC 183). Private, full-subnet administration is handled via encrypted Tailscale mesh (tailscaled on docker-edge advertising 192.168.0.0/24). See Remote Access Architecture.
  4. Active-Backup Network Bonding: Nodes utilize bond0 (Active-Backup Mode 1) bridging nic1 (2.5G primary) and nic0 (1G standby) to unified bridge vmbr0. Wake-on-LAN is bound to nic0. See Networking Architecture.
  5. Intel QuickSync Video (QSV): Transcoding permissions in LXC 131 are stabilized across reboots via /etc/udev/rules.d/99-intel-qsv.rules. See GPU Acceleration Guide.
  6. HDD Drive Health & Backup Policy: Array drive sda (ZC1B91PC) operates under a Run-Until-Death policy with an on-site cold spare. PVE-1 also houses a standalone 2TB WD HDD (/dev/sdf) for local Immich cold backups. See ZFS Mass Storage & DR Runbook.
  7. Automated PBS Power Lifecycle & Weekly Backups: PBS is powered on every Sunday at 00:45 AM via Home Assistant Zigbee smart plug AC-restore, receives snapshot backups at 01:00 AM, auto-shuts down ~3h post-boot (with active task delay guards), and has standby power cut automatically. See PBS Backup Runbook.
  8. Immich Offsite Backup (Backrest & Hetzner Storage Box): Backrest v1.13.0 on docker-media (LXC 131) streams encrypted daily Restic snapshots (1 1 * * *) of /mnt/media_root/media/photos to a 1TB Hetzner Storage Box BX11 with Healthchecks.io webhook reporting. See Backrest Immich Guide.
  9. Router Internet Watchdog & Self-Healing: The GL.iNet Flint 2 gateway runs internet_watchdog.sh every 5 min to test DNS/ICMP connectivity, bounce the physical WAN link (15 min), reboot the router (30 min), and send Telegram restoration alerts. See Router Watchdog Guide.
  10. Docker Container Configuration Standard: All Docker workloads across all LXCs/nodes standardize persistent container configurations and application state under /opt/docker_container_config managed via .env variable FOLDER_FOR_CONFIGS=/opt/docker_container_config (and FOLDER_FOR_CONFIGS_ANIME=/opt/docker_container_config/anime on docker-media). See Container Inventory.

📚 Documentation Index (docs/) ​

All granular technical documentation, architecture deep-dives, and operational runbooks are maintained in the docs/ directory:

docs/
├── DOCUMENTATION_GUIDELINES.md                 # Documentation standards & guide for AI agents
│
├── architecture/                               # Cluster topology, hardware & networking
│   ├── cluster-nodes.md                        # Hardware specs, roles, cluster diagrams, HA policy
│   ├── networking.md                           # Dual-NIC bonding (bond0), WoL, .link rules
│   ├── router-wifi-ap-architecture.md          # OpenWrt dual-router & Wi-Fi AP roaming runbook & matrix
│   ├── remote-access.md                        # Ingress architecture: Cloudflare tunnels, Tailscale, Traefik
│   └── router-watchdog.md                      # GL.iNet Flint 2 router & internet watchdog self-healing script
│
├── storage/                                    # ZFS storage pools, mappings & permissions
│   ├── storage-map.md                          # Datacenter storage map, physical disks & cold backup drive
│   ├── zfs-tank.md                             # ZFS tank RAIDZ1 pool, datasets, drive health policy
│   └── lxc-permissions.md                      # Unprivileged LXC UID/GID shift & POSIX ACLs
│
├── services/                                   # Guest containers, VMs & hardware pass-through
│   ├── container-inventory.md                  # Complete inventory of all LXCs, VMs & Docker stacks
│   ├── backrest-immich.md                      # Backrest (Restic) offsite photo backups to Hetzner Storage Box
│   └── gpu-acceleration.md                     # Intel QuickSync (QSV) pass-through & udev rules
│
└── runbooks/                                   # Step-by-step procedures & disaster recovery
    ├── universal-media-storage.md              # Cross-node LXC storage migration runbook
    ├── drive-replacement-dr.md                 # ZFS drive failure & replacement procedure
    ├── pbs-backup-sync.md                      # Proxmox Backup Server integration & sync architectures
    ├── cold-backup-pipeline.md                 # Automated Plug-and-Run offline cold backup to 2TB WD drive
    ├── pve1_reimage_zfs_runbook.md             # PVE-1 ZFS root re-imaging & recovery runbook
    └── ksm-power-optimization.md               # Proxmox KSM power optimization & C-state tuning

🤖 AI & Maintainer Documentation Guidelines ​

When adding or updating services, storage, networking, or runbooks, adhere to the standard documented in docs/DOCUMENTATION_GUIDELINES.md:


📁 Repository Directory Structure ​

text
homelab-ops/
├── README.md                 # Executive homelab overview & documentation hub (this file)
├── docs/                     # Detailed architectural docs, runbooks, and guides
│   ├── DOCUMENTATION_GUIDELINES.md
│   ├── architecture/
│   ├── storage/
│   ├── services/
│   └── runbooks/
├── docker-admin/             # Core admin services (N8N, Uptime Kuma)
├── docker-media/             # Media server stack configurations (Sonarr, Radarr, etc.)
│   ├── radarr/
│   └── sonarr/
├── homepage-docker/          # Dashboard / homepage configuration (Homepage, Glance, Dozzle)
├── komodo-child-test/        # Komodo manager / child node deployment tests
└── scripts/                  # Automation, monitoring, power management, and maintenance
    ├── cold_backup.sh        # Automated Plug-and-Run offline cold backup script
    ├── internet_watchdog.sh  # Multi-tier self-healing watchdog for GL.iNet Flint 2 router
    ├── systemd/              # Custom systemd unit files (cold-backup.service)
    └── udev/                 # Custom udev rules (99-cold-backup.rules)

Last audited & updated against live cluster state: September 2026.

Authoritative operational repository and DR hub.