open-webui-embeddings · git:20260729.457cf82 · 2026-07-29 · sha256 7bedbe2ff0d2d860
open-webui-embeddings git:20260729.457cf82A
Immutable. This exact content is served forever at /api/v1/blob/7bedbe2ff0d2d860.
---
name: open-webui-embeddings
description: |-
Wire HuggingFace embedding + reranker models (BGE-M3, BGE-Reranker-v2-m3, etc.) into Open WebUI's RAG pipeline via LiteLLM proxying HuggingFace Text Embeddings Inference (TEI). Covers the exact wire shapes Open WebUI sends (URL auto-append on embed but NOT rerank; payload + response shapes for both modes), the LiteLLM-TEI gotchas (encoding_format=null trap, HF-driver task_type misdetection, openai vs huggingface driver tradeoffs), TEI config cliffs (max-client-batch-size 422 under hybrid search, max-batch-tokens AS the auto-truncate boundary, arch-specific Docker images), and the end-to-end production config. BGE-M3 + BGE-Reranker-v2-m3 are worked examples; patterns generalise to any TEI encoder.
when_to_use: |-
Trigger on "open-webui rag", "RAG_OPENAI_API_BASE_URL", "RAG_EXTERNAL_RERANKER_URL", "ExternalReranker", "litellm tei rerank", "encoding_format null", "TEI 422 batch size", "BGE-M3 deployment", "open-webui rerank 404", "Cohere rerank shape", "/v1/rerank vs /rerank", "max-batch-tokens trained ceiling", "tei docker image hangs", "RAG_EMBEDDING_CONCURRENT_REQUESTS". NOT for Open WebUI chat-completion routing, multimodal, or UI/auth; or non-TEI embedding backends (sentence-transformers in-process, Ollama, vLLM, Infinity).
---
# Open WebUI embeddings + reranking — operator reference
Target: operators wiring Open WebUI's RAG pipeline to HuggingFace Text Embeddings Inference (TEI) via LiteLLM. Three hops, each with its own wire-shape quirks. Most failure modes silently degrade to "answer quality dropped" rather than visible errors — this skill is a triage for catching them at config-time.
Verified against **v0.11.0** source (2026-07-29). The embed and rerank code paths are unchanged from 0.10.2 — 0.11.0's retrieval churn landed almost entirely in `retrieval/web/` (web search), not the embedding core — so every wire shape and env default below still holds. One adjacent 0.11.0 change does affect ingestion verification: see §Verifying ingestion below.
**Siblings in the `open-webui` plugin.** Setting these values *through the REST
API* rather than the settings UI — and the knowledge/RAG endpoints that ingest
documents — is the **`open-webui-api`** skill. Running more than one Open WebUI
replica, where RAG requests and their WebSocket streams must survive hitting a
different pod, is **`open-webui-valkey-websocket`**.
## The architecture in 30 seconds
```
Open WebUI → LiteLLM proxy → TEI (GPU)
└ embed: openai-driver → /v1/embeddings
└ rerank: huggingface-driver → /rerank (Cohere↔TEI translation)
```
Why proxy through LiteLLM rather than point Open WebUI at TEI directly?
- **Embed:** TEI exposes `/v1/embeddings` natively (OpenAI-compat) — direct path works. LiteLLM adds: virtual-key auth, per-model rate limits, request logging, optional caching.
- **Rerank:** TEI's native `/rerank` is `{query, texts}` → `[{index, score}]`. Open WebUI's `ExternalReranker` sends Cohere shape `{query, documents, top_n}` → `{results: [{index, relevance_score}]}`. **Direct path fails with HTTP 422** — wire shapes do not match. LiteLLM's HuggingFace rerank handler translates between the two.
Skipping LiteLLM is therefore feasible only for embed; rerank requires either LiteLLM (or another Cohere↔TEI shim) unless Open WebUI itself is patched.
## Wire shapes (exact)
### Embed — Open WebUI code path
`backend/open_webui/retrieval/utils.py:862` (`generate_openai_batch_embeddings`, v0.11.0):
```http
POST {RAG_OPENAI_API_BASE_URL}/embeddings ← URL is auto-appended
Authorization: Bearer {RAG_OPENAI_API_KEY}
Content-Type: application/json
{"input": ["text1", "text2", ...], "model": "{RAG_EMBEDDING_MODEL}"}
```
Response parsed as `data["data"][i]["embedding"]` (OpenAI shape).
Async fan-out (`get_embedding_function` at `utils.py:1090`, batching + `asyncio.gather` at `utils.py:1139-1156`, v0.11.0): chunks bundled into batches of `RAG_EMBEDDING_BATCH_SIZE` (default `1`); all batches dispatched concurrently via `asyncio.gather`, wrapped in an `asyncio.Semaphore` only when `RAG_EMBEDDING_CONCURRENT_REQUESTS` is non-zero (default `0` = unlimited). A 100-chunk file at default config fires **100 concurrent single-chunk requests**.
### Rerank — Open WebUI code path
`backend/open_webui/retrieval/models/external.py:13` (`ExternalReranker`, `predict` at line 26; v0.11.0, unchanged since 0.10.2):
```http
POST {RAG_EXTERNAL_RERANKER_URL} ← URL is exact, NOT appended
Authorization: Bearer {RAG_EXTERNAL_RERANKER_API_KEY}
Content-Type: application/json
{"model": "{RAG_RERANKING_MODEL}", "query": "...",
"documents": ["doc1", "doc2", ...], "top_n": <len(documents)>}
```
Response parsed: `data["results"]` sorted by `index`, extracts `relevance_score`. Cohere shape, strict.
Failure handling: `requests.post()` exception or non-2xx → `predict()` returns `None` → retrieval silently downgrades to **un-reranked hybrid order**. No user-visible error in Open WebUI. **Always alert on rerank-side 4xx in TEI/LiteLLM logs.**
## Open WebUI environment variables
| Variable | Mode | Notes |
|---|---|---|
| `RAG_EMBEDDING_ENGINE` | embed | Set to `openai`. Works for OpenAI, LiteLLM, TEI direct, vLLM direct — anything OpenAI-compat. |
| `RAG_OPENAI_API_BASE_URL` | embed | Open WebUI appends `/embeddings`. Set to `http://litellm:4000/v1` (proxy) or `http://tei:8080/v1` (direct). |
| `RAG_OPENAI_API_KEY` | embed | Bearer token. TEI ignores; LiteLLM enforces virtual key. |
| `RAG_EMBEDDING_MODEL` | embed | Sent in payload as `model`. Must match LiteLLM's `model_name` exactly (case-sensitive, full HF path). |
| `RAG_EMBEDDING_BATCH_SIZE` | embed | Texts per HTTP request. Default `1`. Bumping to `32` reduces per-request overhead during indexing. Legacy alias `RAG_EMBEDDING_OPENAI_BATCH_SIZE` still honoured as a fallback (`config.py:1001-1002`). |
| `RAG_EMBEDDING_CONCURRENT_REQUESTS` | embed | Concurrency cap. Default `0` = unlimited (`asyncio.gather` without semaphore). Set to a bounded number (4-8) to avoid bursting TEI. |
| `RAG_EMBEDDING_PREFIX_FIELD_NAME` | embed | Extra field name for prefix-needing models (e.g. `prompt` for EmbeddingGemma). Leave unset for BGE-M3 — its query/passage symmetry is built into the model. |
| `RAG_EMBEDDING_QUERY_PREFIX` / `RAG_EMBEDDING_CONTENT_PREFIX` | embed | Prefix strings (paired with the field name above). Unused for BGE-M3. |
| `RAG_RERANKING_ENGINE` | rerank | Set to `external` for Cohere-shape endpoints. |
| `RAG_EXTERNAL_RERANKER_URL` | rerank | **Full URL including path** (no auto-append). E.g. `http://litellm:4000/v1/rerank`. |
| `RAG_EXTERNAL_RERANKER_API_KEY` | rerank | Bearer token. |
| `RAG_RERANKING_MODEL` | rerank | Sent in payload as `model`. Match LiteLLM's `model_name`. |
| `RAG_EXTERNAL_RERANKER_TIMEOUT` | rerank | Seconds. Bump for very large `Top_K × Hybrid Search` candidate pools. |
## Triage table
| Symptom | First check | Where |
|---|---|---|
| Embed returns 400 with `encoding_format: expected value` | Add `encoding_format: float` to the LiteLLM litellm_params | `references/gotchas.md` §1 |
| Embed returns 422 with `inputs: data did not match...` | Switch to openai driver — HF driver's task_type detection failed | `references/gotchas.md` §2 |
| Rerank returns 422 with `batch size N > maximum allowed batch size M` | Bump TEI `--max-client-batch-size` | `references/gotchas.md` §3 |
| Rerank returns 404 on `POST /v1` | Open WebUI rerank URL needs full path including `/v1/rerank` | `references/gotchas.md` §7 |
| Open WebUI "Retrieved 1 source" but answer quality dropped | Rerank is silently 4xx — check TEI/LiteLLM logs | `references/gotchas.md` §3 |
| TEI pod hangs at "Starting FlashBert model" | Wrong arch image — match GPU compute capability | `references/gotchas.md` §5 |
| TEI returns 429 during knowledge-base upload | Open WebUI concurrency too high; cap `RAG_EMBEDDING_CONCURRENT_REQUESTS` | `references/gotchas.md` §6 |
| Reranker quality degraded since recent config change | `--max-batch-tokens` past trained ceiling lets long inputs through | `references/gotchas.md` §4 |
| `vector_db` directory growing fast | ChromaDB is fine to ~1 GB; past that switch to pgvector halfvec | `references/gotchas.md` §8 |
## Verifying ingestion (changed in 0.11.0)
The obvious way to confirm a knowledge base actually extracted text — list its files and look at the content — **stopped working in 0.11.0**. `Knowledges.get_file_metadatas_by_id` became a column-only `SELECT File.id, File.hash, File.meta, File.created_at, File.updated_at` (`models/knowledge.py:684-694`), deliberately excluding `File.data`, which is where the extracted text lives. Its own docstring says so.
Consequence: `GET /api/v1/knowledge/{id}` and the KB detail view now return file entries with **no extracted content**, so an empty-looking listing no longer distinguishes "extraction failed" from "extraction succeeded, field not selected". To verify extraction actually produced text, fetch the individual file (`GET /api/v1/files/{id}`) rather than reading the collection listing. Knowledge list items gained a `file_count` in exchange.
This is a pure API-shape change — embedding and retrieval quality are unaffected. It matters only for ingestion-verification scripts and health checks.
## Reference index
- **`references/gotchas.md`** — nine gotchas with HTTP error strings, root causes, and fixes. Load when triage table points here.
- **`references/end-to-end-config.md`** — full working LiteLLM + Open WebUI + TEI config (BGE-M3 + BGE-Reranker-v2-m3 worked example). Load when bootstrapping a new deployment.
- **`references/performance.md`** — quality verification (cross-engine numerical-identity check) + throughput baseline. Load for sizing or post-deployment health checks.
- **`references/sources.md`** — authoritative source files and PR/issue URLs underlying every claim. Load to verify a specific claim or run `freshen` mode.