[RTX 5090] → [Ollama] → [Open WebUI / API REST] → [Utilisateur / n8n / RAG pgvector]
During an earlier diagnostic (investigating abrupt shutdowns under certain agentic workloads), the RTX 5090 was capped by software at 400 W, instead of the default 600 W. The ComfyUI and Ollama tests of September 27–28, 2026 were then run at 600 W, by explicit decision. No lasting stability gain nor performance comparison between these caps has been established.
The cap affects neither VRAM nor model size. The 400 W procedure below keeps a record of the diagnostic; it no longer describes the current state.
nvidia-smi \ --query-gpu=power.limit,power.default_limit \ --format=csv,noheader # → 600.00 W, 600.00 W
Diagnostic value (historical): sudo nvidia-smi --power-limit=400 (not persistent). Return to the factory cap: sudo nvidia-smi --power-limit=600.
| Family | Models | Usage |
|---|---|---|
| Meta | Llama 3.x | High-quality general-purpose LLM |
| Mistral AI | Mistral 7B/22B | Efficient LLM, good performance/size ratio |
| Gemma 2 | Compact and fast LLM | |
| Microsoft | Phi-3 | Small-footprint LLM |
| NousResearch | Hermes | Advanced AI agent — 🔵 validated (design platform) |
| Tool | Role | Version | Port |
|---|---|---|---|
| Ollama | Local LLM inference server | 0.32.6 | 11434 |
| Open WebUI | Chat interface (ChatGPT-style) + RAG | 0.11.0 | 3000 |
| n8n | AI workflow orchestration | — | 5678 |
# Lister les modèles installés
ollama list
# Télécharger un modèle
ollama pull mistral
ollama pull llama3
# Inférence directe
ollama run mistral
# API REST
curl http://localhost:11434/api/generate \\
-d '{"model": "mistral", "prompt": "Résume en 3 points :", "stream": false}'- Local LLM hosting (private AI infrastructure, no cloud)
- Comparison and selection of LLM models based on needs
- Ollama configuration (models, API, inference parameters)
- Open WebUI deployment as the user interface
- LLM integration into pipelines via REST API
- AI agent architecture with Hermes + n8n
The Linux AI node (VM210) now runs ComfyUI 0.36.0 and Ollama on the same RTX 5090 (32 GB). The RealVisXL V5.0 model (photorealistic SDXL derivative) produced portraits at 512×512 and 1,024×1,024. Coexistence with qwen3:8b was observed, then unloading models without stopping the servers was verified. Automatic request sharing still needs to be built (manual switching for now).
| Test | Observation | Scope |
|---|---|---|
| 512×512 portrait, 8 steps | 8.03 s (including initial load) | Functional validation |
| 1,024×1,024 portrait, 30 steps | Image obtained (DPM++ SDE Karras) | Duration not reliably attributed |
| Qwen3 8B + RealVisXL | 19,001 MiB of VRAM (sample) | Qwen 100% GPU, 40,960 context |
| ComfyUI unload | 7,376 → 742 MiB · HTTP 200 | Server kept active |
| Reload + 8-step generation | 7.40 s · back to 7,376 MiB | Full cycle without restart |
| Qwen3 8B unload (ComfyUI stopped) | Back to 2 MiB | Ollama service still available |
MODEL_FILE='<chemin-du-checkpoint>' EXPECTED_SHA256='<empreinte-publiee-du-checkpoint>' printf '%s %s\n' "$EXPECTED_SHA256" "$MODEL_FILE" | sha256sum --check nvidia-smi nvidia-smi --query-gpu=power.limit,power.default_limit --format=csv,noheader
The expected hash must come from the model's trusted repository, not from the downloaded file alone.
nvidia-smi \ --query-gpu=timestamp,temperature.gpu,power.draw,memory.used,utilization.gpu \ --format=csv,noheader --loop=1
Lessons and limitations
- A graph mixing RealVisXL with SD3/AuraFlow nodes was discarded; the adapted SDXL graph then worked
- A blocked cancellation required a targeted restart, without establishing the exact cause of the block
- Anatomical realism, poses and quality at scale remain to be evaluated — one successful portrait is not enough to validate them
- Agentic/n8n integrations, shared admission control and the restart-after-reboot policy remain open work
Detailed results, the photographic need and the full access procedure on the AI Image page. Container inspection and cache fix on the Docker page.