
What is retrieval augmented generation
Retrieval-augmented generation (RAG) is a way to give a generative AI model relevant information from external sources before it answers. A system retrieves useful passages and supplies them as context. The model generates a response from that context, but the evidence and the answer still need checking.
Those sources might be visitor guides, product manuals or an organization’s documents. Context is the information available to the model while it produces an answer. RAG gives generative AI selected evidence to work with, including material that may be private or recently updated.
Supplying that evidence does not require retraining the generator or imply that it has never encountered similar information. If generation and context are new to you, What Is GPT explains those foundations. Here, we will follow the evidence from documents to an answer, using one fictional museum example.
How a RAG system uses documents

Prepare and index
For a searchable document collection, preparation happens before questions arrive. Extract readable text and retain useful headings. Chunking means splitting material into smaller passages that can be retrieved separately; a passage should keep enough surrounding information to make sense.
An index organizes information for search. Preserve each passage’s source, section, effective date, version and access rules. Together, these details help establish provenance: where the evidence came from. Refresh changed content and permissions. Some systems retrieve directly from another service instead of maintaining their own index.
Retrieve passages
At question time, the retriever searches for relevant evidence. A question with two parts may need passages from different documents. Retrieve only material the current user may access. Microsoft's RAG overview describes both content preparation and the importance of authorized, relevant retrieval.
Supply context and generate
The application commonly builds a prompt containing the question, selected passages and separate instructions for answering. The model then generates a response using that working context. Include source identifiers so claims can be traced back to evidence.
Keep document text separate from instructions governing the assistant. Retrieved text is evidence to inspect, even when it contains sentences that look like commands.
Verify or abstain
Check whether each factual claim follows from the supplied passages and whether the sources apply to the question. More context is not automatically better; irrelevant material can obscure useful evidence. When support is missing, answer only the supported parts or abstain, meaning decline to make an unsupported claim. Verification can involve automated checks and human review; neither should be assumed infallible.
A museum question from sources to answer
Read the documents
This is a fictional example. The museum documents and responses are invented for teaching, not visitor advice for a real museum. Our collection contains three short records:
- A, Visitor guide. Current, effective October 1, 2026. Hours section: “Saturday opening hours are 10 am to 4 pm. Last entry is 3:30 pm.”
- B, Access note. Current, effective October 1, 2026. Entrance section: “The east entrance has a ramp.”
- C, Visitor guide. Archived June 2026 version, explicitly superseded by A. Hours section: “Saturday opening hours are 10 am to 5 pm. Last entry is 4:30 pm.”
Preparation keeps the A, B and C labels attached to their passages. C remains clearly marked as archived. If vector retrieval is used, an embedding model also creates numerical representations for similarity search. This preparation stores searchable evidence; it does not put the museum’s opening hours into the generator’s learned parameters, or weights.
Follow the question
The visitor asks: “On Saturdays, what time is last entry, and which entrance has a ramp?”
The system searches for both last entry and ramp access. In this successful walkthrough, it retrieves A’s Hours passage and B’s Entrance passage. For a current visit, version metadata excludes C because A explicitly replaces it.
The context supplied to the generator contains the question and the two selected excerpts, labeled A and B. Separate instructions tell it to answer only supported parts, cite the source identifiers and acknowledge missing evidence. The full archive need not be placed in the prompt.
An illustrative answer is: “Saturday last entry is 3:30 pm [A, Hours]. The east entrance has a ramp [B, Entrance].”

Now inspect the evidence trail. A supports the time; B supports the entrance. Neither excerpt establishes parking availability, ticket prices or café hours, so the answer should not add them. In a real application, each citation should open the original document or exact passage. Here, the labels refer only to our fictional records.
Keyword vector and hybrid retrieval

Keyword retrieval matches search terms to document text. It can be useful for a precise phrase such as “last entry” or a distinctive product code.
Vector retrieval compares embeddings: numerical representations of questions and passages. It can find related wording even when exact words differ. For example, “step-free entrance” might help find a passage about a ramp, but the match still needs checking. Word Embeddings Explained explores the underlying representations.
Hybrid retrieval combines keyword and vector signals. Microsoft's query documentation distinguishes these three approaches. Embeddings are common in RAG, but they are not compulsory, and a dedicated vector database is not a universal requirement.
All three methods select candidate evidence. A high similarity score does not establish truth, currentness or permission to disclose a document. Knowledge Representation offers optional background on other ways to organize knowledge.
How RAG differs from related approaches
Search alone returns results for a person to inspect. RAG adds generation using retrieved evidence, although a modern search product may itself include a RAG-style answer feature. Pasting a document into a prompt supplies context; by itself, it does not include an automated retrieval stage.
Fine-tuning changes a model’s learned parameters, often called weights. Retrieving passages at answer time does not itself change those weights. A system can use both techniques. The original RAG paper by Lewis and colleagues studied a trainable retriever-generator setup; its training recipe is not required for every RAG application. For parameter adaptation, see Transfer Learning Explained.
Memory stores or reuses information across interactions. Tools give an application capabilities, such as searching a collection or checking a calendar. Either can incorporate retrieval, but having memory or tools does not automatically make a system RAG. Chatbots Explained covers those roles in conversation systems.
Why RAG answers can still be wrong
Missing evidence and old versions
Change the museum walkthrough to expose three failures:
- Retrieval miss: only A is returned. Give the supported last-entry time and say the retrieved material does not establish which entrance has a ramp. Search again if appropriate; do not guess “west.”
- Wrong version: C is returned and the answer says “4:30 pm.” The citation supports that quotation, but the guidance is outdated. Check effective dates and supersession. Unresolved conflicts require a caveat or abstention, rather than blindly choosing whichever document looks newest.
- Insufficient evidence: the visitor asks whether the café is open. None of A–C answers this. Say, “These documents do not state the café opening hours.” Missing evidence does not mean the café is closed.

Check claims and citations
Even after good retrieval, a generator can add unsupported details or misread a passage. Sources can also contain mistakes. Finding relevant evidence, generating a supported answer and being correct about the world are separate achievements.
A citation’s presence proves little on its own. Does it resolve to a source? Does that passage support this exact claim? Is the source reliable and applicable? Gao and colleagues' citation research evaluates answer correctness and citation quality separately. RAG can help reduce unsupported answers, but it does not eliminate hallucinations: false or fabricated content presented as factual.

Use documents safely and test the system
Enforce access rules before private text reaches the model. Minimize the data sent, and check how the search service and model provider transmit, store and log it. Keep sources, indexes and permissions up to date; adding retrieval does not make stale material fresh.
Documents can contain malicious instructions. A page might tell the assistant to ignore its rules or expose private information. This is indirect prompt injection, a risk described by OWASP. Treat retrieved instructions as untrusted, limit connected actions and enforce access outside the model. Prompt wording alone cannot guarantee protection.
Test with answerable questions, missing evidence, conflicting versions and permission-restricted documents. Check whether the needed passages were found, every claim is supported, citations open correctly and partial answers or abstentions are appropriate. Include user review and track response time and cost alongside quality. The museum scenarios illustrate these checks; they are not measured results from a deployed system.
Frequently asked questions
Does RAG train the model on my documents?
Ordinary retrieval at answer time does not update the generator’s weights. A service’s separate data-retention or training policies still matter, so check them before providing private material.
Does RAG require embeddings or a vector database?
No. Keyword retrieval can supply evidence too. Vector and hybrid approaches use embeddings, but the storage and search design depends on the application.
How is RAG different from a search engine?
Search alone finds results. RAG also generates an answer using retrieved information. Search engines with AI answers may combine both functions.
Is pasting a document into a prompt RAG?
It provides context, but pasting alone lacks automated retrieval. Some document-upload products search the uploaded material behind the scenes; their implementation determines whether retrieval is involved.
Can RAG and fine-tuning be used together?
Yes. Fine-tuning can adapt model behavior while retrieval supplies relevant evidence at answer time. They address different parts of the system and still need evaluation.
Does RAG guarantee accurate answers?
No. Retrieval can miss evidence, sources can be wrong or outdated, and generation can produce unsupported claims. Important answers need checks beyond fluent wording.
Why can a cited answer still be wrong?
The cited passage may not support the claim, may be unreliable or may apply to an old version. Inspect the evidence, not just the source marker.
Can RAG use private or recently updated documents?
Yes, with an authorized connection and appropriate access controls. Updated information must actually reach the retrieval system; freshness and privacy are implementation responsibilities.
What should happen when the documents do not answer the question?
State what is missing, give any supported partial answer and search again when appropriate. Do not turn absent evidence into a confident factual claim.
Key takeaway and next step
RAG connects a question to external evidence and then to a generated answer. Follow that trail: retrieve the right passages, check their status and verify each claim. Next, read Model Evaluation Metrics Explained to connect evaluation choices to the task, alongside the RAG-specific checks above.