Without retrieval: the model answers from itself
Ask a plain language model a question and it generates a plausible continuation based on everything it absorbed in training. Sometimes that's right. Sometimes it's a confident, fluent invention. There is no source, so there is nothing to check, and crucially the two cases *look identical*.
With retrieval: the model answers from documents
- 1
Your document is chopped into passages
A PDF becomes a few hundred chunks of a few hundred words each, each one keeping a note of the page it came from. That page reference is what a citation is made of later.
- 2
Each passage is turned into a vector
An embedding — a long list of numbers positioning the passage in a space where 'similar meaning' means 'physically close'. This is how the system can find the passage about enzyme inhibition without you using the word 'enzyme'.
- 3
Your question becomes a vector too
And the system fetches the passages nearest to it. This is retrieval: not a keyword search, a meaning search.
- 4
The model is asked about those passages only
Instead of 'what's the answer', it's asked 'what do these passages say'. Now the answer has a source, and the source has a page number.
The part that most tools get wrong
Retrieval always returns something. Ask about a topic your document doesn't cover and it will still hand back the three least-irrelevant passages — because 'nearest' is defined even when nothing is near. The model then receives three useless passages and a question, and answers anyway, blending them with its own prior knowledge in the same confident voice.
So RAG on its own is not enough. It needs a relevance floor: if nothing retrieved clears a similarity threshold, generation must not happen at all. That is the difference between a tool that says "not in your material" and one that quietly makes something up with a citation stapled to it. The longer argument is here.