⚠️ Advanced skill: GPU Passthrough VFIO is one of the most complex and sought-after achievements in the portfolio. It allows the RTX 5090 to be physically assigned to a VM, with performance close to a native installation.
Why an RTX 5090 for local AI? With its CUDA and Tensor cores and its 32 GB of GDDR7 memory, it strongly accelerates the parallel computation needed for language model inference, embeddings and image generation — while keeping the models and processed data local.
Why passthrough despite its constraints? PCIe passthrough with VFIO directly assigns the card to a VM: the guest system, CUDA and containers use the GPU with near-native performance and clear workload isolation. In return, the assignment is exclusive (no simultaneous use by several VMs) and requires rigorous preparation of the host, boot process and recovery procedures.
⚡ Power cap — diagnostic measurement (past)

A temporary 400 W software cap (versus 600 W by default on the observed card) was used during an electrical diagnostic. The ComfyUI and Ollama tests of September 27–28, 2026 were then run at 600 W, by explicit decision. No lasting stability gain nor performance comparison between these caps has been established.

reproducible public check
nvidia-smi \
  --query-gpu=power.limit,power.default_limit \
  --format=csv,noheader

The 400 W diagnostic value is applied with sudo nvidia-smi --power-limit=400 — a setting that is not persistent after a reboot or driver reload. Detailed commands, telemetry and the return to the factory cap are documented on the AI & LLM (VM210) page.

🔧 Technical stack
TechnologyRole
VFIOLinux PCIe isolation framework
IOMMU / AMD-ViMemory management unit for I/O
PCI PassthroughDirect GPU → VM assignment
Q35QEMU chipset required for modern PCIe
OVMFUEFI required for NVIDIA passthrough
📁 Modified system files
FileChange
/etc/default/grubamd_iommu=on iommu=pt
/etc/modprobe.d/vfio.confbind VFIO to GPU IDs
/etc/modulesvfio vfio_iommu_type1 vfio_pci
/etc/initramfs-tools/Module integration at boot
/etc/pve/qemu-server/*.confProxmox VM config
🐛 Layered diagnosis

The setup required distinguishing several diagnostic levels, so that each effect could be attributed to a single parameter:

  • firmware and IOMMU
  • host driver and PCIe groups
  • reservation of the GPU and its audio function to vfio-pci
  • VM configuration
  • guest driver and CUDA
  • containerized runtime (GPU access all the way into Docker)

Before any firmware update, the settings required for IOMMU and VFIO are saved and a rollback is prepared; after any change, IOMMU groups, GPU reservation and VM startup are revalidated. A separate virtual console is kept for recovery.

Not shown here: no one-off workaround (OVMF patch, hypervisor masking, particular IOMMU mapping…) is displayed until an exact public proof is available. Validating VFIO, the guest driver and CUDA does not allow them to be inferred.
🔄 GPU exclusivity — VM status and startup order

Run on the Proxmox host, this command detects the RTX 5090 (without displaying its address), identifies the VMs configured to use it, and outputs a sanitized label — distinguishing the VM's current state from its auto-start policy, and checking that only one passthrough VM is active at a time.

bash — GPU exclusivity check
gpu_bdf=$(lspci -D | awk '/NVIDIA.*RTX 5090/ {print $1; exit}')
[ -n "$gpu_bdf" ] || { echo 'RTX 5090 non détectée'; exit 1; }

gpu_slot=${gpu_bdf%.*}
gpu_slot_short=${gpu_slot#0000:}
alternative=0
running_gpu_vms=0

for vmid in $(qm list | awk 'NR > 1 {print $1}'); do
    config=$(qm config "$vmid")

    if grep -Eq "^hostpci[0-9]+:.*(${gpu_slot}|${gpu_slot_short})([.,]|$)" \
        <<< "$config"; then
        if [ "$vmid" = 210 ]; then
            label="IA-Core"
        else
            alternative=$((alternative + 1))
            label="GPU-alternative-${alternative}"
        fi

        status=$(qm status "$vmid" | awk '{print $2}')
        onboot=$(awk -F': ' '/^onboot:/ {print $2}' <<< "$config")
        startup=$(awk -F': ' '/^startup:/ {print $2}' <<< "$config")

        [ -z "$onboot" ] && onboot=0
        [ -z "$startup" ] && startup="-"
        [ "$status" = running ] && running_gpu_vms=$((running_gpu_vms + 1))

        printf '%-20s status=%-8s onboot=%-2s startup=%s\n' \
            "$label" "$status" "$onboot" "$startup"
    fi
done

if [ "$running_gpu_vms" -le 1 ]; then
    printf 'GPU passthrough actif : %s VM — exclusivité OK\n' "$running_gpu_vms"
else
    printf 'ALERTE : %s VM GPU actives simultanément\n' "$running_gpu_vms"
    exit 2
fi
sanitized excerpt — exclusivity verified
IA-Core              status=running  onboot=1  startup=order=20,up=30,down=60
GPU-alternative-1    status=stopped  onboot=0  startup=-
GPU-alternative-2    status=stopped  onboot=0  startup=-
…                    status=stopped  onboot=0  startup=-
GPU passthrough actif : 1 VM — exclusivité OK
Reading: only one passthrough VM (IA-Core / VM210) uses the RTX 5090; alternative GPU environments stay out of auto-start. On host resume, VM215 (PostgreSQL/pgvector) starts first, then VM210 (GPU workload + Ollama / OpenWebUI / RAG) — with no contention for the GPU.
🔧 BIOS baseline & firmware caution

A reference BIOS baseline separates values actually observed from improvements merely considered. Before any firmware update, the parameters needed for IOMMU and VFIO are backed up and a rollback is prepared; after the change, the IOMMU groups, the GPU reservation and VM startup are revalidated. Tests remain one change at a time, so that each effect can be clearly attributed to the modified parameter.

Structuring constraint: the RTX 5090 is an exclusive physical resource — only one passthrough VM can use it at a time. The architecture therefore explicitly documents the simultaneous-startup incompatibilities, and a separate virtual console is kept for recovery.

Validation: the NVIDIA driver and CUDA were validated in an Ubuntu VM, then GPU access was verified down into Docker containers (via the NVIDIA Container Toolkit).

Related pages

References & Sources

CategoryResourceURL
Official documentationProxmox VE — PCI(e) Passthroughpve.proxmox.com/wiki/PCI_Passthrough
Official documentationLinux Kernel — VFIOdocs.kernel.org/driver-api/vfio
HardwareNVIDIA GeForce RTX 5090 · plateforme AMD (IOMMU / AMD-Vi)—
LicenseVFIO (noyau Linux) — GPL · pilotes NVIDIA — propriétaires—
Content of this pageShared under CC BY-SA 4.0creativecommons.org/licenses/by-sa/4.0