A temporary 400 W software cap (versus 600 W by default on the observed card) was used during an electrical diagnostic. The ComfyUI and Ollama tests of September 27–28, 2026 were then run at 600 W, by explicit decision. No lasting stability gain nor performance comparison between these caps has been established.
nvidia-smi \ --query-gpu=power.limit,power.default_limit \ --format=csv,noheader
The 400 W diagnostic value is applied with sudo nvidia-smi --power-limit=400 — a setting that is not persistent after a reboot or driver reload. Detailed commands, telemetry and the return to the factory cap are documented on the AI & LLM (VM210) page.
| Technology | Role |
|---|---|
| VFIO | Linux PCIe isolation framework |
| IOMMU / AMD-Vi | Memory management unit for I/O |
| PCI Passthrough | Direct GPU → VM assignment |
| Q35 | QEMU chipset required for modern PCIe |
| OVMF | UEFI required for NVIDIA passthrough |
| File | Change |
|---|---|
| /etc/default/grub | amd_iommu=on iommu=pt |
| /etc/modprobe.d/vfio.conf | bind VFIO to GPU IDs |
| /etc/modules | vfio vfio_iommu_type1 vfio_pci |
| /etc/initramfs-tools/ | Module integration at boot |
| /etc/pve/qemu-server/*.conf | Proxmox VM config |
The setup required distinguishing several diagnostic levels, so that each effect could be attributed to a single parameter:
- firmware and IOMMU
- host driver and PCIe groups
- reservation of the GPU and its audio function to
vfio-pci - VM configuration
- guest driver and CUDA
- containerized runtime (GPU access all the way into Docker)
Before any firmware update, the settings required for IOMMU and VFIO are saved and a rollback is prepared; after any change, IOMMU groups, GPU reservation and VM startup are revalidated. A separate virtual console is kept for recovery.
Run on the Proxmox host, this command detects the RTX 5090 (without displaying its address), identifies the VMs configured to use it, and outputs a sanitized label — distinguishing the VM's current state from its auto-start policy, and checking that only one passthrough VM is active at a time.
gpu_bdf=$(lspci -D | awk '/NVIDIA.*RTX 5090/ {print $1; exit}')
[ -n "$gpu_bdf" ] || { echo 'RTX 5090 non détectée'; exit 1; }
gpu_slot=${gpu_bdf%.*}
gpu_slot_short=${gpu_slot#0000:}
alternative=0
running_gpu_vms=0
for vmid in $(qm list | awk 'NR > 1 {print $1}'); do
config=$(qm config "$vmid")
if grep -Eq "^hostpci[0-9]+:.*(${gpu_slot}|${gpu_slot_short})([.,]|$)" \
<<< "$config"; then
if [ "$vmid" = 210 ]; then
label="IA-Core"
else
alternative=$((alternative + 1))
label="GPU-alternative-${alternative}"
fi
status=$(qm status "$vmid" | awk '{print $2}')
onboot=$(awk -F': ' '/^onboot:/ {print $2}' <<< "$config")
startup=$(awk -F': ' '/^startup:/ {print $2}' <<< "$config")
[ -z "$onboot" ] && onboot=0
[ -z "$startup" ] && startup="-"
[ "$status" = running ] && running_gpu_vms=$((running_gpu_vms + 1))
printf '%-20s status=%-8s onboot=%-2s startup=%s\n' \
"$label" "$status" "$onboot" "$startup"
fi
done
if [ "$running_gpu_vms" -le 1 ]; then
printf 'GPU passthrough actif : %s VM — exclusivité OK\n' "$running_gpu_vms"
else
printf 'ALERTE : %s VM GPU actives simultanément\n' "$running_gpu_vms"
exit 2
fiIA-Core status=running onboot=1 startup=order=20,up=30,down=60 GPU-alternative-1 status=stopped onboot=0 startup=- GPU-alternative-2 status=stopped onboot=0 startup=- … status=stopped onboot=0 startup=- GPU passthrough actif : 1 VM — exclusivité OK
A reference BIOS baseline separates values actually observed from improvements merely considered. Before any firmware update, the parameters needed for IOMMU and VFIO are backed up and a rollback is prepared; after the change, the IOMMU groups, the GPU reservation and VM startup are revalidated. Tests remain one change at a time, so that each effect can be clearly attributed to the modified parameter.
Validation: the NVIDIA driver and CUDA were validated in an Ubuntu VM, then GPU access was verified down into Docker containers (via the NVIDIA Container Toolkit).