Architecture : [RTX 5090] → [Ollama] → [Open WebUI / API REST] → [Utilisateur / n8n / RAG pgvector]
Ollama, Open WebUI and the RAG run on the main AI node VM210 (Ubuntu Server 24.04 LTS · Docker Engine + NVIDIA Container Toolkit · RTX 5090), with the RAG relying on PostgreSQL/pgvector hosted on VM215. VM260 (Windows) is an alternative machine for PoC / quick validation, where components are proven before being scheduled on the final infrastructure. Since the RTX 5090 is exclusive, only one GPU node uses it at a time.
⚡ RTX 5090 power cap — past diagnostic

During an earlier diagnostic (investigating abrupt shutdowns under certain agentic workloads), the RTX 5090 was capped by software at 400 W, instead of the default 600 W. The ComfyUI and Ollama tests of September 27–28, 2026 were then run at 600 W, by explicit decision. No lasting stability gain nor performance comparison between these caps has been established.

The cap affects neither VRAM nor model size. The 400 W procedure below keeps a record of the diagnostic; it no longer describes the current state.

check of the active cap / default cap
nvidia-smi \
  --query-gpu=power.limit,power.default_limit \
  --format=csv,noheader
# → 600.00 W, 600.00 W

Diagnostic value (historical): sudo nvidia-smi --power-limit=400 (not persistent). Return to the factory cap: sudo nvidia-smi --power-limit=600.

🤖 Deployed LLM models
FamilyModelsUsage
MetaLlama 3.xHigh-quality general-purpose LLM
Mistral AIMistral 7B/22BEfficient LLM, good performance/size ratio
GoogleGemma 2Compact and fast LLM
MicrosoftPhi-3Small-footprint LLM
NousResearchHermesAdvanced AI agent — 🔵 validated (design platform)
🖥️ Interfaces & Tools
ToolRoleVersionPort
OllamaLocal LLM inference server0.32.611434
Open WebUIChat interface (ChatGPT-style) + RAG0.11.03000
n8nAI workflow orchestration—5678
💻 Ollama commands
bash
# Lister les modèles installés
ollama list

# Télécharger un modèle
ollama pull mistral
ollama pull llama3

# Inférence directe
ollama run mistral

# API REST
curl http://localhost:11434/api/generate \\
  -d '{"model": "mistral", "prompt": "Résume en 3 points :", "stream": false}'
🎯 Skills demonstrated
  • Local LLM hosting (private AI infrastructure, no cloud)
  • Comparison and selection of LLM models based on needs
  • Ollama configuration (models, API, inference parameters)
  • Open WebUI deployment as the user interface
  • LLM integration into pipelines via REST API
  • AI agent architecture with Hermes + n8n
🖼️ Local image generation — ComfyUI & GPU sharing

The Linux AI node (VM210) now runs ComfyUI 0.36.0 and Ollama on the same RTX 5090 (32 GB). The RealVisXL V5.0 model (photorealistic SDXL derivative) produced portraits at 512×512 and 1,024×1,024. Coexistence with qwen3:8b was observed, then unloading models without stopping the servers was verified. Automatic request sharing still needs to be built (manual switching for now).

session observations — RTX 5090 VRAM (not a benchmark)
TestObservationScope
512×512 portrait, 8 steps8.03 s (including initial load)Functional validation
1,024×1,024 portrait, 30 stepsImage obtained (DPM++ SDE Karras)Duration not reliably attributed
Qwen3 8B + RealVisXL19,001 MiB of VRAM (sample)Qwen 100% GPU, 40,960 context
ComfyUI unload7,376 → 742 MiB · HTTP 200Server kept active
Reload + 8-step generation7.40 s · back to 7,376 MiBFull cycle without restart
Qwen3 8B unload (ComfyUI stopped)Back to 2 MiBOllama service still available
These are session observations, with no statistical protocol or latency guarantee. A coexistence reading shows 539 W · 53 °C · 83% GPU (600 W configured cap) — not to be presented as constant consumption or a measured peak. Larger LLMs, longer contexts or other workflows may exceed the available memory. RealVisXL V5.0 is distributed under the Open RAIL++ license: available weights do not mean an absence of license restrictions.
Reproducible check — checkpoint and GPU verification
MODEL_FILE='<chemin-du-checkpoint>'
EXPECTED_SHA256='<empreinte-publiee-du-checkpoint>'
printf '%s  %s\n' "$EXPECTED_SHA256" "$MODEL_FILE" | sha256sum --check

nvidia-smi
nvidia-smi --query-gpu=power.limit,power.default_limit --format=csv,noheader

The expected hash must come from the model's trusted repository, not from the downloaded file alone.

Reproducible check — GPU telemetry during a generation (Ctrl+C to stop)
nvidia-smi \
  --query-gpu=timestamp,temperature.gpu,power.draw,memory.used,utilization.gpu \
  --format=csv,noheader --loop=1

Lessons and limitations

  • A graph mixing RealVisXL with SD3/AuraFlow nodes was discarded; the adapted SDXL graph then worked
  • A blocked cancellation required a targeted restart, without establishing the exact cause of the block
  • Anatomical realism, poses and quality at scale remain to be evaluated — one successful portrait is not enough to validate them
  • Agentic/n8n integrations, shared admission control and the restart-after-reboot policy remain open work

Detailed results, the photographic need and the full access procedure on the AI Image page. Container inspection and cache fix on the Docker page.

Related pages

References & Sources

CategoryResourceURL
Official documentationOllamagithub.com/ollama/ollama
Official documentationOpen WebUIdocs.openwebui.com
Version usedOllama 0.32.6 · Open WebUI 0.11.0 · Ubuntu Server 24.04 LTS (nœud VM210)—
LicenseOllama — MIT · Open WebUI — BSD-3-Clause—
Content of this pageShared under CC BY-SA 4.0creativecommons.org/licenses/by-sa/4.0