Host: the AI containers (Ollama, Open WebUI, RAG) run on the main node VM210 (Ubuntu Server 24.04 · Docker Engine + NVIDIA Container Toolkit — CUDA access to the RTX 5090 is validated down into the containers). The vector backend PostgreSQL / pgvector is on VM215 (no GPU); Qdrant is kept as a fallback solution. VM260 (Windows) serves as a PoC / validation bench (e.g. ComfyUI via Docker Desktop).
📦 Deployed services
ServiceImagePortStatus
Ollamaollama/ollama · 0.32.6 · VM21011434🟢 Operational
ComfyUI0.36.0 · hardened local image · VM210 (Linux)8188🔵 Validated (local generation)
Open WebUIghcr.io/open-webui/open-webui · 0.11.0 · VM2103000🟢 Operational
PostgreSQL + pgvectordeployed on VM215 · 18.4 / 0.8.65432🟢 Operational
Qdrantqdrant/qdrant · 1.18.36333🟡 Fallback
n8nn8nio/n8n:latest · VM3005678🟢 Operational
NPM (Nginx Proxy Manager)jc21/nginx-proxy-manager:latest443 · 81⚪ Unsourced
Portainerportainer/portainer-ce:latest9443⚪ Unsourced
Whisper APIonerahmet/openai-whisper-asr-webservice · VM2109000🔵 Validated · planned
🎯 Skills demonstrated
  • Docker Engine + Docker Compose installation
  • Writing docker-compose.yml files
  • Persistent volume management (ZFS mount)
  • Docker network management (bridge, host)
  • Application-level isolation of AI services
  • Environment variables via .env files
📄 docker-compose.yml
yaml
services:
  open-webui:
    image: ghcr.io/open-webui/open-webui:main   # 0.11.0
    container_name: open-webui
    ports:
      - '3000:8080'
    volumes:
      - open-webui-data:/app/backend/data
    environment:
      - OLLAMA_BASE_URL=http://host-gateway:11434
      # Backend vectoriel : pgvector hébergé sur VM215 (sans GPU)
      - VECTOR_DB=pgvector
      - PGVECTOR_DB_URL=postgresql://<role>:<secret>@<hote-vm215>/rag
      - RAG_EMBEDDING_ENGINE=ollama
      - RAG_EMBEDDING_MODEL=bge-m3
    extra_hosts:
      - 'host-gateway:host-gateway'
    restart: unless-stopped

  n8n:
    image: n8nio/n8n:latest
    container_name: n8n
    ports:
      - '5678:5678'
    volumes:
      - n8n-data:/home/node/.n8n
    restart: unless-stopped

  # Solution de repli — conservée tant que le retrait n'est pas clôturé
  qdrant:
    image: qdrant/qdrant:latest                 # 1.18.3
    container_name: qdrant
    ports:
      - '6333:6333'
    volumes:
      - qdrant-data:/qdrant/storage
    restart: unless-stopped

volumes:
  open-webui-data:
  n8n-data:
  qdrant-data:

The n8n service above is a simplified example (local volume, no external database) meant to illustrate the Compose structure. The actual deployment on VM300 relies on a dedicated external PostgreSQL database; see the Automation (n8n) page for the real architecture.

🖼️ Hardened ComfyUI container (VM210)

ComfyUI 0.36.0 runs in a locally built image: Python base pinned by digest, sources pinned by commit, dependencies locked with their hashes (wheels downloaded then installed offline). Validated environment: Python 3.12.14, PyTorch 2.11.0, CUDA 13.0.

  • Unprivileged execution, read-only root filesystem
  • Models mounted read-only; dedicated folders for inputs, outputs, state and caches
  • Custom nodes and API disabled (this is not a general outbound network block)
  • Interface published only on the VM's local loopback; access via SSH tunnel, no Internet publication
  • Checkpoint verified by SHA-256 hash (from the trusted repository) before use

The earlier Windows PoC (VM260) stays separate from this Linux deployment. ComfyUI/Ollama GPU sharing is a manual switch, not yet an automatic scheduler. Details and measurements on the AI & LLM page.

🔎 Container inspection and cache fix

Command actually used (sanitized) — status checks run against the ComfyUI container and the Ollama container: container status, last log lines, resident Ollama models, then HTTP checks of the ComfyUI local API. Container and host names replaced with examples.

bash
COMFY_CONTAINER='<conteneur-comfyui>'
OLLAMA_CONTAINER='<conteneur-ollama>'
COMFY_URL='http://127.0.0.1:8188'

docker inspect "$COMFY_CONTAINER" --format '{{.State.Status}}'
docker logs --tail 40 "$COMFY_CONTAINER"
docker exec "$OLLAMA_CONTAINER" ollama ps
curl --fail --max-time 10 -sS -o /dev/null \
  -w 'HTTP=%{http_code}\n' "$COMFY_URL/"
curl --fail --max-time 10 -sS "$COMFY_URL/object_info/CheckpointLoaderSimple"
curl --fail --max-time 10 -sS "$COMFY_URL/queue"

HTTP 200 confirms the server responded; the checkpoint's presence in the loader confirms it was detected, not yet a successful inference. ollama ps lists resident models, without guaranteeing there are no requests in flight. Details on memory release (/free) and GPU telemetry on the AI & LLM page.

Command actually used (sanitized) — a startup once failed in getpass.getuser(): the host account's UID did not exist in the container's passwd file. The following variables kept the unprivileged execution model, without rebuilding the image:

yaml
environment:
  USER: comfyui
  TORCHINDUCTOR_CACHE_DIR: /state/cache/torchinductor

/state/cache here denotes an example mount writable by the container's UID. The USER variable does not create an account and grants no privilege.

Related pages

References & Sources

CategoryResourceURL
Official documentationDocker Documentationdocs.docker.com
Official documentationNVIDIA Container Toolkitdocs.nvidia.com — container-toolkit
Version usedOllama 0.32.6 · Open WebUI 0.11.0 · PostgreSQL 18.4 · pgvector 0.8.6 · Qdrant 1.18.3 (repli)—
LicenseDocker Engine (Moby) — Apache 2.0github.com/moby/moby — LICENSE
Content of this pageShared under CC BY-SA 4.0creativecommons.org/licenses/by-sa/4.0