Learn how Retrieval-Augmented Generation combines search and generation.
Complete 12 of 24 practices (50%) and enter your name to unlock the Certificate of Participation.
RAG is a system pattern: retrieve relevant external information at request time, place selected evidence into the model context, and ask the LLM to answer using that evidence.
Instead of asking an analyst to remember the whole filing cabinet, RAG lets the analyst search first and answer with the relevant documents on the desk.
hits = retrieve(question, top_k=5)
context = format_context(hits)
answer = llm(
question=question,
context=context
)RAG does not make the model omniscient.
RAG quality starts before the user asks a question.
If important evidence is missing from the index, the generator cannot recover it.
record = {
"source_id": "policy_042",
"section": "Leave",
"text": chunk_text,
"effective_date": "2026-01-01",
"vector": embed(chunk_text)
}The user's question may be vague, multi-part or phrased differently from the source documents.
A retrieval layer may transform the request into a better search representation without changing the user's intent.
search_query = rewrite(
user_question,
preserve_intent=True
)
hits = retrieve(
search_query,
filter=user_scope,
top_k=10
)Query transformation is useful only when it preserves user intent.
Retrieved chunks are not automatically a good prompt.
Three diverse, authoritative chunks may be more useful than ten duplicates.
SYSTEM = (
"Answer using only EVIDENCE. "
"If evidence is insufficient, say so. "
"Cite source IDs for factual claims."
)
prompt = SYSTEM + format(hits)A strong RAG prompt tells the model what evidence is authoritative and what to do when evidence is missing or conflicting.
A fluent unsupported answer is worse than a concise statement that the available evidence is insufficient.
According to HR Policy §4, eligible employees may submit Form HR-17 within the stated process. Source: policy_042.
Insufficient evidenceThe retrieved sources do not establish that requirement.
Grounding means connecting claims to evidence.
Once the basic pipeline works, teams may add more advanced retrieval stages.
Add a stage when evaluation shows a specific problem that the stage can improve.
candidates = retrieve(question, top_k=20)
ranked = reranker.rank(
query=question,
documents=candidates
)
context = ranked[:5]RAG can fail at different layers, so evaluate the layers separately.
Diagnose whether retrieval or generation failed.
for case in eval_set:
hits = retrieve(case.question)
retrieval_score = score_retrieval(
hits, case.expected_sources
)
answer = generate(case.question, hits)
answer_score = score_answer(
answer, case.expected_facts
)Production RAG combines data pipelines, vector indexes, prompts, models, filters and evaluation.
For a production answer, you should be able to identify the components that produced it.
trace = {
"index_version": "hr_2026_09",
"retriever": "hybrid_v3",
"prompt_version": "rag_answer_v5",
"model_version": MODEL_VERSION,
"source_ids": [h.id for h in hits],
"latency_ms": latency_ms
}The next training, #10 Evals, goes deeper into systematic measurement.
Open each item after answering it in your own words.
Retrieval selects relevant external evidence; generation uses selected evidence as context to produce an answer.
Retrieved chunks must be selected, deduplicated, labeled and formatted so useful evidence fits the context and remains traceable.
Follow an explicit fallback policy, state the limitation, and avoid presenting unsupported claims as established facts.
Because a RAG failure may come from different layers.
Versioned components, evaluation, access control, observability, freshness monitoring and safe fallback behavior.
| Layer | Purpose |
|---|---|
| Sources | Maintain approved, current knowledge |
| Ingest | Parse, chunk, label and version source content |
| Index | Create embeddings and searchable vector / hybrid indexes |
| Retrieve | Find eligible evidence for the user's request |
| Assemble | Select, deduplicate, label and format context |
| Generate | Answer from evidence with clear fallback rules |
| Evaluate | Measure retrieval, faithfulness, relevance, citations and system behavior |
| Observe | Track versions, source IDs, latency, cost, freshness and failures |
Complete at least 12 of the 24 practice cases (50%) and enter your name.
RAG behavior depends on source quality, retrieval design, model behavior and workflow constraints.