🖥️ Hardware components
| Component | Detail | Role |
|---|---|---|
| CPU | AMD Ryzen 9 9950X3D — 16 cores / 32 threads | AI inference · KVM virtualization |
| GPU | NVIDIA GeForce RTX 5090 | LLM inference · Image generation |
| RAM | 192 GB DDR5 (~186 GiB usable) | Multiple VM memory · large models |
| Storage | 2× NVMe SSD in ZFS mirror (~1.75 TiB usable) | Virtual disks · AI datasets · resilience |
| Hypervisor | Proxmox VE 9.2.2 | Virtualization · GPU allocation |
| Network | Ethernet | Proxmox bridge · LAN access |
⚙️ BIOS/UEFI — attested value
| Setting | State | Purpose |
|---|---|---|
| IOMMU (AMD) | ✅ Enabled and validated | VFIO GPU passthrough |
The other BIOS values (Above 4G Decoding, Re-Size BAR, Secure Boot…) are not published: only IOMMU activation is attested (see GPU / VFIO). A working passthrough does not replace a dated BIOS report.
🎯 Skills demonstrated
- High-performance local AI server architecture
- Hardware selection for GPU inference (VRAM, bandwidth)
- CPU / RAM sizing for a multi-VM hypervisor
- BIOS/UEFI configuration for virtualization and PCIe passthrough
- High-performance NVMe storage management with ZFS
🧩 Role distribution — virtual machines
| VM | Role | Base | Key components |
|---|---|---|---|
| VM200 | Reverse proxy — Web entry point / HTTPS publishing | Ubuntu Server 24.04 LTS · no GPU | Nginx · Certbot (Let's Encrypt) · Fail2ban · local firewall |
| VM205 | File service — Nextcloud | Ubuntu Server 24.04 LTS · no GPU | Nextcloud + MariaDB (Docker containers) · auto-start |
| VM210 | AI node — GPU workloads | Ubuntu Server 24.04 LTS · RTX 5090 (passthrough) | Docker + NVIDIA Container Toolkit · Ollama · Open WebUI |
| VM215 | Shared services — no GPU | Ubuntu Server 24.04 LTS · dedicated ext4 disk | PostgreSQL 18.4 · pgvector 0.8.6 (RAG vector backend) |
| VM300 | AI consumer — automation | Ubuntu Server 24.04 LTS · no GPU | n8n (workflow orchestration) · auto-start |
Principle: the GPU is assigned directly to the VMs that need it; GPU workloads (VM210) and shared CPU services (VM215) are kept separate to limit dependencies and ease backups.
Industrialization: the VMs are derived from generic and CUDA templates, customized via Cloud-Init; rebuilding them follows a reversible procedure (a control clone is validated before any replacement).
Limitations: a diagnostic 400 W software cap was used previously (versus 600 W by default); the latest ComfyUI/Ollama tests were run at 600 W, with no guarantee of long-term stability nor an established performance comparison (procedure on the AI & LLM page). ZFS snapshots also ease technical rollbacks but do not replace an independent backup — that policy is still being defined.
Industrialization: the VMs are derived from generic and CUDA templates, customized via Cloud-Init; rebuilding them follows a reversible procedure (a control clone is validated before any replacement).
Limitations: a diagnostic 400 W software cap was used previously (versus 600 W by default); the latest ComfyUI/Ollama tests were run at 600 W, with no guarantee of long-term stability nor an established performance comparison (procedure on the AI & LLM page). ZFS snapshots also ease technical rollbacks but do not replace an independent backup — that policy is still being defined.