Why Proxmox for this project? The hypervisor is the base that makes the platform scalable and agile: each role (AI node, shared services, PoC bench) is added, isolated or evolves without touching the others. It brings a real time saving — in case of corruption, an entire server is recloned in a few minutes from a template or a snapshot, instead of reinstalling everything. And it extends the possibilities: Linux / Windows coexistence on the same machine, including one or more Windows clients to consume the AI services, with GPU allocation on demand.
📦 Version & features
ItemValue
Proxmox VE9.2.2
PVE kernel6.17
VirtualizationKVM + QEMU
UEFI VMOVMF
Legacy BIOSSeaBIOS
Paravirt driversVirtIO
📋 Available templates
NameOSRoleStatus
WIN11-CUDAWindows 11 ProDesktop + GPU passthrough🟢 Operational
Ubuntu Desktop-CUDAUbuntu Desktop 24.04GPU passthrough / CUDA / Docker🟢 Operational
WIN11-baseW11-debloat-sysprepAI dev/test workstation🟢 Operational
Ubuntu Server-CUDAUbuntu Server 24.04AI-CORE servers + GPU passthrough + CUDA / Docker🟢 Operational
Ubuntu ServerUbuntu Server 24.04Functional servers with no GPU need🟢 Operational
🔄 VM status & startup order

This command produces directly from Proxmox the list of VMs with their current state, their auto-start policy (onboot) and their resume order (startup) — without confusing the instantaneous state (running / stopped) with the persistent configuration.

bash — status + startup order
for vmid in $(qm list | awk 'NR > 1 {print $1}'); do
    config=$(qm config "$vmid")

    name=$(awk -F': ' '/^name:/ {print $2}' <<< "$config")
    status=$(qm status "$vmid" | awk '{print $2}')
    onboot=$(awk -F': ' '/^onboot:/ {print $2}' <<< "$config")
    startup=$(awk -F': ' '/^startup:/ {print $2}' <<< "$config")

    [ -z "$onboot" ] && onboot=0
    [ -z "$startup" ] && startup="-"

    printf '%-6s %-30s status=%-8s onboot=%-2s startup=%s\n' \
        "$vmid" "$name" "$status" "$onboot" "$startup"
done
sanitized excerpt — resume order: VM215 (10) → VM210 (20) → VM300 n8n (30) → VM205 (40)
205    205-Ubuntu-Nextcloud           onboot=1  startup=order=40,up=30,down=60
210    210-IA-CORE-CUDA               onboot=1  startup=order=20,up=30,down=60
215    215-IA-Services                onboot=1  startup=order=10,up=60,down=60
260    260-W11-CUDA-Ollama            onboot=0  startup=-
300    300-n8n                        onboot=1  startup=order=30,up=30,down=60
…      environnements à la demande    onboot=0  startup=-
Reading: VMs with onboot=1 resume in increasing startup order — VM215 (services · PostgreSQL/pgvector) → VM210 (AI node) → VM300 (n8n) → VM205 (Nextcloud) — so that dependencies are ready before their consumers. Alternative GPU environments stay out of auto-start. The exhaustive inventory remains in the internal operations documentation.
🎯 Skills demonstrated
  • Proxmox VE administration (web UI + CLI)
  • Creation, configuration and management of KVM/QEMU VMs
  • ZFS storage management via Proxmox
  • Creation of reusable system templates
  • Proxmox snapshots (vzdump) — not yet an independent backup
  • VM cloning for rapid deployment
  • Hypervisor troubleshooting (logs, VNC console)
🧪 Lessons learned — template rebuild (VM130)

Cloning an old template revealed a residual hostname. The clone was fixed and checked, then the generic template was rebuilt from an official ISO to explicitly control its content. The build VM was sanitized before conversion: machine identity, hostname, SSH keys, network, Cloud-Init state and logs were cleaned.

A full validation clone confirmed Cloud-Init initialization, generation of a clean identity, networking, DNS, SSH, the guest agent, and persistence after reboot. The old template was only replaced after this check; the ZFS origins were verified before deleting the candidate that had become redundant, to preserve VM independence.

♻️ Template lifecycle & types

A template is not worth much for its cloning speed but for the mastery of what it transmits (system, identity, network, access, initialization). Building, validating and replacing are three separate operations: the old template stays available until its replacement and its validation clone have met every criterion.

template lifecycle
Source système vérifiée
        │
        ▼
VM temporaire de construction
        │ installation et configuration du socle
        ▼
Neutralisation
        │ identités, réseau, secrets et journaux
        ▼
Template temporaire
        │
        ▼
Clone de validation
        │ contrôles fonctionnels et d'indépendance
        ▼
Template validé
        │
        ├── clone complet → VM applicative
        └── template dérivé → socle spécialisé
Base typeExpected contentExcluded elements
Generic serverminimal system, Cloud-Init, SSH, guest agentapplication, secret, business configuration
Specialized servergeneric base + justified common dependenciesdata specific to one VM
CUDA / GPUvalidated GPU drivers and toolsapplication unrelated to the GPU
Windowsprepared system, VirtIO drivers, generalization mechanismidentity and account specific to the source VM
A specialized template must not make the generic one dependent on its components; applications normally stay installed in the clones.
🐛 Lessons learned — PCIe diagnosis

An onboard PCIe device intermittently disappeared. The diagnosis distinguished a real absence on the bus from a simple driver fault, by successively checking the interface, the modules, PCIe enumeration, the parent port and the rescan. A full power cycle restored the device; recurrence is being monitored before considering a conditional firmware update.

Related pages

References & Sources

CategoryResourceURL
Official documentationProxmox VE Administration Guidepve.proxmox.com/pve-docs
GitHub repositoryproxmox/pve-managergithub.com/proxmox/pve-manager
Version usedProxmox VE 9.2.2 — KVM/QEMU · OVMF · VirtIO—
LicenseAGPLv3 (Community Edition)gnu.org/licenses/agpl-3.0
Content of this pageShared under CC BY-SA 4.0creativecommons.org/licenses/by-sa/4.0