Retrieval-Augmented Generation at Scale: Lessons from Production
Abstract
Retrieval-Augmented Generation (RAG) is the dominant pattern for grounding LLMs in enterprise knowledge bases. This paper documents lessons from operating RAG at production scale.
Chunking Strategy
Fixed-size chunking is insufficient. We recommend semantic chunking with overlap, tuned per document type:
- Technical docs: 512 tokens, 64 token overlap
- Legal/compliance: sentence-boundary chunking
- Code: function-level chunks with import context
Hybrid Search
Pure vector search misses exact matches (SKUs, error codes). Hybrid search combining BM25 + dense retrieval improved recall by 23% in our benchmarks.
Evaluation
Pre-deploy eval suites with golden Q&A pairs are non-negotiable. We run 200+ test cases on every index update.
Conclusion
RAG at scale is an engineering discipline, chunking, retrieval, and evaluation must be treated as first-class infrastructure.