
How RAG Pipelines Really Work: Chunking, Embeddings, and Reranking
Retrieval-Augmented Generation sounds simple: ingest documents, convert them into vectors, find the most relevant sections for a query, pass them to the language model. In practice, most implementations fail on three underestimated problems: bad chunking, poor embeddings, and missing reranking.
Chunking: The Underestimated Foundation
Most teams use fixed chunk sizes: 512 tokens, 1024 tokens. That works for homogeneous documents, but it fails against reality: product manuals have different structures than support articles. API documentation is organized differently than FAQs. Fixed chunk sizes tear semantic units apart — an answer gets split across two chunks, and retrieval only finds half of it.
The better strategy is context-aware chunking: detect the document type, respect natural boundaries (paragraphs, headings, code blocks), and overlap chunks for context continuity. It costs more preprocessing, but it pays off massively in retrieval quality.
Embeddings: Not All Vectors Are Equal
The market for embedding models is hard to navigate: OpenAI, Cohere, Voyage, BGE, E5 — every model has its own strengths. The decisive factor is not the MTEB benchmark score, but the semantic density in the target language. German texts need embeddings that understand German compound words and sentence structures. An English model that does not recognize "Kundenservice" and "Service für Kunden" as equivalent leaves gaps even on simple queries.
In practice, that means evaluating embedding quality regularly with domain-specific test queries, not just trusting generic benchmarks. Measure retrieval recall on real customer questions, not on synthetic datasets.
Reranking: The Underestimated Quality Lever
Most RAG systems take the top-k results from vector similarity search and pass them straight to the language model. The problem: vector similarity measures semantic proximity, not relevance. Two texts can be semantically close (both are about returns) yet differ completely in relevance to the actual query ("How long will my refund take?").
Reranking adds a second evaluation step: a specialized model (a cross-encoder) compares the query directly against every candidate chunk and scores actual relevance. It is more compute-intensive than pure vector search, but the quality improvement is significant — especially for complex or ambiguous queries.
A System Emerges Only Through Integration
Chunking, embeddings, and reranking are not independent optimization problems. Better chunking changes which embeddings get produced. Better embeddings shift which chunks make it into reranking. And reranking compensates for weaknesses in both upstream steps.
Teams that want to seriously improve their RAG pipeline should not optimize a single parameter in isolation. They should treat the entire chain as a system — and start where the biggest quality loss occurs. You won't find that in a benchmark report. You find it by running real customer queries through the pipeline and measuring where answer quality breaks down.
Author
Sammy
Sammy leitet den Vertrieb bei Reaktly und hilft Unternehmen dabei, das Potenzial von KI im Kundenservice zu erschließen. Er findet für jedes Team die passende Lösung.
Reaktly
Ready to transform your customer service?
See how Reaktly combines AI agents, knowledge management, and omnichannel support in one platform.