The usual RAG recipe is embed the documents, store the vectors, fetch the nearest ones and hand them to the LLM. It makes a convincing demo, but it skips decades of what information retrieval learned the hard way. Words change meaning with context, exact keyword matches get lost in embeddings, chunk size changes what you find, and the LLM can still ignore or misread the context you give it.
Takeaways
- Real search systems work in two stages: a fast, cheap filter to get candidates, then a slower ranker to order them. One embedding lookup does neither job well.
- Bigger chunks do not give the LLM better answers. Retrieval quality sets the ceiling for answer quality.
- Treat RAG as an ongoing project: build the evaluation pipeline first, curate the data, and measure against business numbers, not only model scores.
Read the full article →
Have a system that needs a second opinion?