My Homelab Two Years Later, Every Upgrade and What It Cost
A two-year homelab evolution from a 7,440 dual-server rack with 73 containers shows that financial break-even hinges on a single $80/month security monitoring substitution, while the real return was infrastructure engineering skills that transferred directly to professional work.
Key Points
- Evolved from a 7,440 hardware spend.
- Rack averages 430W metered (~0.15/kWh); CPUs average 13% busy, 10G links peak at ~12% of one link — significant over-provisioning.
- Night mode automation saves only ~5W (not the claimed 175W) because stopped containers were idle; CPU governor drift on prxbox2 went undetected because hosts aren’t in OpenTofu.
- Every container now defined in OpenTofu with drift detection; 185 health checks catch failures but missed a frozen power sensor for 3 weeks — check 186 will monitor sensor staleness.
- Break-even ~4.6 years (131/mo net savings) but leans 40%+ on replacing an $80/mo security monitoring plan; without it, break-even exceeds 11 years.
- Real ROI was professional: running the full stack (network, hosts, deploys, monitoring, power) built infrastructure engineering skills that transfer directly to day-job work.
- [AI Synthesis] The project illustrates how hobbyist infrastructure often over-provisions compute while under-investing in observability and network bottlenecks (NAS on 1G uplink).
- [AI Synthesis] DDR4 ECC RAM prices rose ~5× due to AI cluster demand, making memory the new hidden cost in used enterprise hardware — buy what you’ll use upfront.
Hardware Evolution
- Phase 1: Lenovo ThinkServer (Xeon E3-1226 v3, 4C/32GB) ran Plex, *arr stack, Pi-Hole for 6 months until Plex transcoding and Frigate AI detection saturated CPU.
- Phase 2: Two Dell PowerEdge R730s (dual Xeon E5-2698 v4, 40C/80T each, 128GB/host) from Server Design Lab (~$1,850 each configured) for PCIe GPU slots, ECC RAM, iDRAC, dual PSUs.
- Phase 3: Added Quadro RTX 4000 (8GB) on prxbox1 for always-on Frigate/Whisper/Tdarr; RTX A4000 (16GB) on prxbox2 for bursty Plex/Immich/Tdarr/Ollama LLMs.
- Rack: StarTech 25U, MikroTik CRS317 (10G), UniFi US-24 (1G), Synology DS418play (48TB), 2× CyberPower UPS, UniFi Cloud Gateway Ultra + 3× U7 Pro APs.
GPU & PCIe Constraints
- 2U servers require single-slot blower-style workstation GPUs; gaming cards exceed height, connector clearance, and airflow limits — RTX 4090 does not fit at all.
- Both pro cards draw >75W, requiring Dell GPU power cables off the riser; Quadro RTX 4000 averages ~41W, RTX A4000 ~25W.
- VRAM usage: prxbox1 steady ~4GB (Frigate + Whisper + Tdarr); prxbox2 bursts to 15GB for 30B coding model but averages <1GB — models unload when idle.
Networking & VLANs
- Dual 10G LACP bonds (Intel X710) per host to MikroTik — 20Gbps theoretical, but 30-day peaks ~1.2Gbps (12% of one 10G link); NAS on 1G uplink is the bottleneck.
- Four VLANs via UniFi Cloud Gateway Ultra: Default LAN (trusted), Guest (internet + speakers), IoT (Home Assistant + music + internet), Cameras (Frigate + HA + DNS) — IoT/Guest/Cameras blocked from Main LAN.
- mDNS reflection across VLANs is overly broad; 11 devices on IoT VLAN, older smart devices still on main LAN behind MikroTik firewall list.
Container Workloads & OpenTofu
- 73 LXC containers (45 on prxbox1, 28 on prxbox2) + Home Assistant VM: media (Plex, Immich, *arr, Tdarr, Audiobookshelf), security (Frigate, Vaultwarden, Authentik, CrowdSec), AI (Ollama, llama.cpp, LiteLLM gateway), monitoring (Prometheus, Grafana, Loki, Alertmanager, Uptime Kuma), productivity (Paperless-NGX + AI tagging, Syncthing).
- All containers defined in OpenTofu with shared
bpg/proxmoxmodule; provision scripts enforce immutability — SSH fixes banned; drift check runstofu planevery 6 hours. - Home Assistant moved into VM on prxbox1 after VLANs enabled isolation; enables 134 automations (night mode, AI alert rewriting, backup recovery, UPS monitoring) but creates single-point-of-failure on prxbox1.
Monitoring & Health Checks
- Prometheus scrapes every 15s (CPU, RAM, disk, network, GPU, VRAM, temps, iDRAC power, rack power); 30-day retention sourced all metrics in article.
- 185 health checks (153 shell + 32 Python): backup verification, SSL expiry, disk thresholds, DNS, container restarts, subtitle generation — each checks one thing, exits 0/1.
- Critical gap: frozen rack power sensor (flat 541W for 3 weeks) undetected; new check will alert on stale sensors. Proxmox hosts not in OpenTofu allowed CPU governor drift.
Power Consumption & Costs
- Metered (7-day smart plug): prxbox1 ~215W, prxbox2 ~191W, both servers ~406W, whole rack ~430W → ~10 kWh/day, ~565/yr at $0.15/kWh.
- Night mode (stop 11 containers 11PM–6AM) saves ~5W, not 175W — containers were idle; overnight backups run in same window. CPU
powersavegovernor drifted on 31/80 threads of prxbox2. - R730s designed for datacenter power/cooling; in Minnesota attic they’re free heat in winter, cooling problem in summer. Mini PCs would run most workloads at fraction of power but can’t hold 16GB GPU or provide iDRAC/ECC.
Break-even Analysis
- Replaced subscriptions: iCloud/Dropbox (13), 1Password→Vaultwarden (20), Ring Protect (80), Audible (21) = ~$187/mo gross.
- Net ~47 power + 7,240 hardware (excl. ThinkServer). Remove $80 security monitoring → break-even >11 years.
- NAS B2 backup and Claude API usage not counted; time investment excluded. Author frames it as a hobby: “Nobody asks a golfer to justify their club membership with a break-even analysis.”
Lessons Learned
- Infrastructure as code from day one — including Proxmox hosts — to prevent config drift (CPU governor, iDRAC network).
- Measure before automating: 5 minutes in Prometheus would have shown night mode saves ~5W, not 175W.
- VLANs from day one: retrofitting segmentation onto running services is a miserable weekend of IP updates.
- Buy for the bottleneck: 10G between servers wasted while NAS stayed on 1G uplink.
- Monitoring before first failure, not after third: uptime checks and backup verification should exist from start.
- Offsite copy before you need it: local backups don’t survive attic fire.
- GPU form factor research: rack servers ≠ desktop cases — measure twice, buy once.
- Budget for RAM upfront but buy only what you’ll use: DDR4 ECC ~5× price spike from AI demand; using 72GB of 256GB.