Why RAG? In OpenWebUI, RAG, or retrieval-augmented generation, lets the user query their own documents in natural language and get a contextualized, updatable answer grounded in the retrieved passages, without retraining the model.
🧭 Why PostgreSQL / pgvector rather than Qdrant?
PostgreSQL with pgvector is chosen as the nominal backend: pgvector is one of the vector backends maintained by the OpenWebUI core team, while Qdrant is a non-core, community-maintained integration there — which increases the risk of breakage or delayed fixes on an OpenWebUI version upgrade. Qdrant remains a valid vector solution, kept as a fallback.
Another factor is storage: Qdrant's active data was not kept on the ZFS share presented to VM210 (via VirtioFS/FUSE) and had to stay on a local ext4 on the VM; PostgreSQL/pgvector has its own dedicated ext4 disk on VM215, better isolated for operations and future logical backup.
Distributed architecture: inference and embedding computation run on the GPU AI node (VM210, RTX 5090). Vector storage is offloaded to a GPU-less services VM (VM215) hosting PostgreSQL and pgvector. This separation limits dependencies and eases backup of the vector backend.
[Documents] → [Chunking + bge-m3 Embeddings · GPU VM210] → [PostgreSQL / pgvector · VM215]
↓
[Question] → [Query embedding] → [HNSW semantic search] → [Local Ollama LLM] → [Cited answer]
↓
[Question] → [Query embedding] → [HNSW semantic search] → [Local Ollama LLM] → [Cited answer]
🔧 Technical stack
| Component | Role | Version |
|---|---|---|
| Open WebUI | Interface + native RAG orchestration | 0.11.0 |
| PostgreSQL | Vector backend (VM215) | 18.4 |
| pgvector | Vector extension · HNSW index | 0.8.6 |
| bge-m3 | Multilingual embeddings (dim 1024) | via Ollama |
| Ollama | LLM inference (VM210 · RTX 5090) | 0.32.6 |
| Qdrant | Fallback vector database | 1.18.3 |
✅ Demonstrated achievements
- Open WebUI configured with pgvector as the vector backend
- bge-m3 embedding model validated at dimension 1024
- HNSW index created on the vectors
- Exact search and citations verified
- Out-of-corpus behavior tested to limit invented answers
- Persistence validated after recreating the interface and restarting both VMs
⚙️ Open WebUI → pgvector configuration (illustrative)
env — pgvector vector backend
# Open WebUI — RAG natif sur backend pgvector (VM215) VECTOR_DB=pgvector PGVECTOR_DB_URL=postgresql://<role>:<secret>@<hote-vm215>/rag # Embeddings servis par Ollama sur le nœud GPU (VM210) RAG_EMBEDDING_ENGINE=ollama RAG_EMBEDDING_MODEL=bge-m3 # dimension 1024
sql — HNSW similarity index (pgvector)
-- Index HNSW pour la recherche vectorielle rapide (distance cosinus) CREATE INDEX ON documents USING hnsw (embedding vector_cosine_ops);
Scope & limits: the validations rely on an entirely synthetic corpus — this is not a business dataset. Qdrant is kept as a fallback solution (not removed) until the phase-out is closed. Final cleanup of the test collections and the closing restore point are still to be done.