Why RAG? In OpenWebUI, RAG, or retrieval-augmented generation, lets the user query their own documents in natural language and get a contextualized, updatable answer grounded in the retrieved passages, without retraining the model.
🧭 Why PostgreSQL / pgvector rather than Qdrant?

PostgreSQL with pgvector is chosen as the nominal backend: pgvector is one of the vector backends maintained by the OpenWebUI core team, while Qdrant is a non-core, community-maintained integration there — which increases the risk of breakage or delayed fixes on an OpenWebUI version upgrade. Qdrant remains a valid vector solution, kept as a fallback.

Another factor is storage: Qdrant's active data was not kept on the ZFS share presented to VM210 (via VirtioFS/FUSE) and had to stay on a local ext4 on the VM; PostgreSQL/pgvector has its own dedicated ext4 disk on VM215, better isolated for operations and future logical backup.

Distributed architecture: inference and embedding computation run on the GPU AI node (VM210, RTX 5090). Vector storage is offloaded to a GPU-less services VM (VM215) hosting PostgreSQL and pgvector. This separation limits dependencies and eases backup of the vector backend.
[Documents] → [Chunking + bge-m3 Embeddings · GPU VM210] → [PostgreSQL / pgvector · VM215]
                                              ↓
[Question] → [Query embedding] → [HNSW semantic search] → [Local Ollama LLM] → [Cited answer]
🔧 Technical stack
ComponentRoleVersion
Open WebUIInterface + native RAG orchestration0.11.0
PostgreSQLVector backend (VM215)18.4
pgvectorVector extension · HNSW index0.8.6
bge-m3Multilingual embeddings (dim 1024)via Ollama
OllamaLLM inference (VM210 · RTX 5090)0.32.6
QdrantFallback vector database1.18.3
✅ Demonstrated achievements
  • Open WebUI configured with pgvector as the vector backend
  • bge-m3 embedding model validated at dimension 1024
  • HNSW index created on the vectors
  • Exact search and citations verified
  • Out-of-corpus behavior tested to limit invented answers
  • Persistence validated after recreating the interface and restarting both VMs
⚙️ Open WebUI → pgvector configuration (illustrative)
env — pgvector vector backend
# Open WebUI — RAG natif sur backend pgvector (VM215)
VECTOR_DB=pgvector
PGVECTOR_DB_URL=postgresql://<role>:<secret>@<hote-vm215>/rag

# Embeddings servis par Ollama sur le nœud GPU (VM210)
RAG_EMBEDDING_ENGINE=ollama
RAG_EMBEDDING_MODEL=bge-m3        # dimension 1024
sql — HNSW similarity index (pgvector)
-- Index HNSW pour la recherche vectorielle rapide (distance cosinus)
CREATE INDEX ON documents USING hnsw (embedding vector_cosine_ops);
Scope & limits: the validations rely on an entirely synthetic corpus — this is not a business dataset. Qdrant is kept as a fallback solution (not removed) until the phase-out is closed. Final cleanup of the test collections and the closing restore point are still to be done.

Related pages

References & Sources

CategoryResourceURL
Official documentationpgvector — extension de recherche vectorielle pour PostgreSQLgithub.com/pgvector/pgvector
Official documentationOpen WebUI — RAG & bases vectoriellesdocs.openwebui.com
Embedding modelBAAI/bge-m3 — multilingue, dimension 1024huggingface.co/BAAI/bge-m3
Version usedOpen WebUI 0.11.0 · PostgreSQL 18.4 · pgvector 0.8.6 · Ollama 0.32.6 · Qdrant 1.18.3 (repli)—
Licensepgvector — PostgreSQL License · PostgreSQL — PostgreSQL Licensepostgresql.org/about/licence
Content of this pageShared under CC BY-SA 4.0creativecommons.org/licenses/by-sa/4.0