Glossary
Retrieval-augmented generation (RAG): what it is
Retrieval-augmented generation (RAG) is a technique in which a system first retrieves relevant information from a set of documents or data and then gives it to a language model to answer from. It lets a model use a company's own, current information without retraining. The term comes from a 2020 research paper by Lewis and colleagues at Facebook AI Research.
Updated · 3 min read
RAG vs fine-tuning vs a longer prompt
| RAG | Fine-tuning | Everything in the prompt | |
|---|---|---|---|
| Adds | Facts, looked up per question | Style, format or a narrow skill | Facts, all at once |
| Data freshness | As current as the index | Fixed at training time | As current as the prompt |
| Scale | Millions of documents | Not a store of facts | Limited by the context window |
| Can cite sources | Yes | No | Yes |
Where RAG goes wrong
- Retrieval returns the wrong passages, so the model answers confidently from the wrong source.
- Documents are split badly, cutting tables and clauses in half.
- Old versions sit in the index next to new ones.
- Nobody tests retrieval on its own, so failures look like model errors.
RAG inside an agent
In production, retrieval is usually one tool an agent uses among several, next to direct database queries and system lookups. Native helped Pierpont Holdings pair a vector store with 8,000+ validated question-to-SQL pairs, so procurement leaders can ask questions of their data in plain English.
Sources
Frequently asked questions
What is RAG in simple terms?
RAG means looking up the relevant parts of your documents first, then asking the AI to answer using only what was found.
Does RAG stop AI hallucinations?
It reduces them by grounding answers in retrieved text, but it does not remove them. Good retrieval, citations and evals on real questions are what make RAG reliable.
Is RAG better than fine-tuning?
For adding facts that change, yes. Fine-tuning suits teaching a model a format or a narrow skill. Many systems use both.

