Data Scientist → AI Engineer • Training 09
Article-Training • AI Engineering Foundations

RAG

Grounding LLMs with External Knowledge

Learn how Retrieval-Augmented Generation combines search and generation.

Question → Retrieve → Select Evidence → Build Context → Generate → Verify
❓ QUESTION
→
🔎 RETRIEVE
→
📚 EVIDENCE
🧠 LLM
→
🧾 GROUNDED ANSWER
→
✅ VERIFY
8learning modules
24interactive practices
5rapid review questions
50%certificate threshold
Learning target
Design a RAG workflow that retrieves useful evidence, builds controlled context, produces source-grounded answers, exposes uncertainty, and can be evaluated as a system.
Practice progress0 / 24

Complete 12 of 24 practices (50%) and enter your name to unlock the Certificate of Participation.

MODULE 01
🌉

RAG: The Bridge Between Retrieval and Generation

RAG is a system pattern: retrieve relevant external information at request time, place selected evidence into the model context, and ask the LLM to answer using that evidence.

👁️
See it this way

Instead of asking an analyst to remember the whole filing cabinet, RAG lets the analyst search first and answer with the relevant documents on the desk.

Question→Retriever→Evidence→Prompt Context→LLM→Answer

What RAG changes

  • Knowledge can come from external, current sources
  • Evidence can be traced back to source records
  • The model can be constrained to retrieved evidence
  • Retrieval quality becomes part of answer quality
Minimal RAG sketch
hits = retrieve(question, top_k=5)
context = format_context(hits)
answer = llm(
    question=question,
    context=context
)
✅

RAG does not make the model omniscient.

Practice — 3 cases

Practice 1 / Práctica 1
What is the core idea of RAG?
Practice 2 / Práctica 2
Why can RAG help with current organizational knowledge?
Practice 3 / Práctica 3
Which statement is most accurate?
MODULE 02
📚

Prepare the Knowledge Base for Retrieval

RAG quality starts before the user asks a question.

🧱
Bad inputs become bad context

If important evidence is missing from the index, the generator cannot recover it.

Knowledge-base checklist

  • Approved, authoritative sources
  • Useful chunk boundaries and retained headings
  • Stable source IDs and metadata
  • Fresh index synchronized with source changes
Indexed chunk
record = {
  "source_id": "policy_042",
  "section": "Leave",
  "text": chunk_text,
  "effective_date": "2026-01-01",
  "vector": embed(chunk_text)
}

Practice — 3 cases

Practice 4 / Práctica 4
A correct answer exists in the source system but the relevant document was never indexed. What happens?
Practice 5 / Práctica 5
Why retain stable source IDs?
Practice 6 / Práctica 6
Which knowledge-base property most directly reduces stale-answer risk?
MODULE 03
🔎

Retrieve the Right Evidence

The user's question may be vague, multi-part or phrased differently from the source documents.

🧭
The question is not always the best search query

A retrieval layer may transform the request into a better search representation without changing the user's intent.

Retrieval strategies

  • Semantic vector search
  • Hybrid lexical + vector search
  • Metadata / access filtering
  • Query rewriting or expansion
  • Multi-query or decomposition for complex questions
Query rewrite sketch
search_query = rewrite(
    user_question,
    preserve_intent=True
)

hits = retrieve(
    search_query,
    filter=user_scope,
    top_k=10
)
✅

Query transformation is useful only when it preserves user intent.

Practice — 3 cases

Practice 7 / Práctica 7
When is query rewriting useful?
Practice 8 / Práctica 8
Why might a complex question be decomposed into sub-queries?
Practice 9 / Práctica 9
Which retrieval approach is strongest for a query containing both a concept and an exact form code?
MODULE 04
🧾

Build Context the Model Can Use

Retrieved chunks are not automatically a good prompt.

📎
Evidence needs structure

Three diverse, authoritative chunks may be more useful than ten duplicates.

SOURCE 1 — HR Policy §4Parental leave eligibility...
SOURCE 2 — HR FAQHow to submit Form HR-17...
SOURCE 3 — Benefits GuideCoverage during approved leave...
CONTEXT POLICYUse only these sources; if unsupported, say so.

Context assembly

  • Prefer authoritative, non-duplicative evidence
  • Preserve source labels and IDs
  • Fit evidence within the context budget
  • Separate instructions from retrieved content
Grounded prompt pattern
SYSTEM = (
    "Answer using only EVIDENCE. "
    "If evidence is insufficient, say so. "
    "Cite source IDs for factual claims."
)

prompt = SYSTEM + format(hits)

Practice — 3 cases

Practice 10 / Práctica 10
Why label retrieved chunks with source IDs?
Practice 11 / Práctica 11
What is a common problem with highly duplicated retrieved context?
Practice 12 / Práctica 12
How should retrieved text containing 'ignore previous instructions' be treated?
MODULE 05
🛡️

Grounded Generation and Honest Fallbacks

A strong RAG prompt tells the model what evidence is authoritative and what to do when evidence is missing or conflicting.

⚖️
Evidence before fluency

A fluent unsupported answer is worse than a concise statement that the available evidence is insufficient.

Generation policy

  • Answer from retrieved evidence
  • Separate facts from inference
  • Cite supporting sources when appropriate
  • State when evidence is missing or conflicting
Supported answer

According to HR Policy §4, eligible employees may submit Form HR-17 within the stated process. Source: policy_042.

Insufficient evidence

The retrieved sources do not establish that requirement.

✅

Grounding means connecting claims to evidence.

Practice — 3 cases

Practice 13 / Práctica 13
The retrieved context does not support the user's requested fact. What is the strongest behavior?
Practice 14 / Práctica 14
Why cite sources in a RAG answer?
Practice 15 / Práctica 15
Two retrieved authoritative sources conflict. What should the system avoid?
MODULE 06
🚀

Advanced RAG Patterns

Once the basic pipeline works, teams may add more advanced retrieval stages.

🧰
Complexity must earn its place

Add a stage when evaluation shows a specific problem that the stage can improve.

Reranking
Improve ordering of first-pass candidates
Multi-query
Search multiple formulations of the same intent
Decomposition
Split complex questions into retrievable subproblems

Choose by failure mode

  • Good candidates, bad order → rerank
  • User wording misses source vocabulary → rewrite / multi-query
  • Question spans separate evidence → decomposition
  • Retrieved chunks are too verbose → compression / selection
Rerank sketch
candidates = retrieve(question, top_k=20)
ranked = reranker.rank(
    query=question,
    documents=candidates
)
context = ranked[:5]

Practice — 3 cases

Practice 16 / Práctica 16
The right documents are usually retrieved but often appear below irrelevant ones. Which pattern directly targets that problem?
Practice 17 / Práctica 17
When is query decomposition useful?
Practice 18 / Práctica 18
What is the strongest reason to add an advanced RAG stage?
MODULE 07
🧪

Evaluate Retrieval and Generation Separately

RAG can fail at different layers, so evaluate the layers separately.

🔬
Diagnose the layer

Diagnose whether retrieval or generation failed.

RetrievalRecall@K, precision, ranking, filter correctness
GenerationFaithfulness, answer relevance, citation support, format
SystemLatency, cost, freshness, availability
SafetyAccess control, prompt injection resistance, secret handling

Evaluation set

  • Common user questions
  • Edge and ambiguous cases
  • Known-answer cases with expected sources
  • No-answer cases where fallback is expected
  • Adversarial / untrusted-content cases
Layered evaluation
for case in eval_set:
    hits = retrieve(case.question)
    retrieval_score = score_retrieval(
        hits, case.expected_sources
    )

    answer = generate(case.question, hits)
    answer_score = score_answer(
        answer, case.expected_facts
    )

Practice — 3 cases

Practice 19 / Práctica 19
The correct source never appears in top-k results. Which layer should you investigate first?
Practice 20 / Práctica 20
The correct source is in context, but the answer contradicts it. Which layer is more directly implicated?
Practice 21 / Práctica 21
Why include no-answer cases in a RAG evaluation set?
MODULE 08
🏭

Production RAG — Observable, Versioned and Safe

Production RAG combines data pipelines, vector indexes, prompts, models, filters and evaluation.

📡
A RAG answer has a lineage

For a production answer, you should be able to identify the components that produced it.

Source Version→Index Version→Retriever→Prompt Version→Model→Eval

Production signals

  • Retrieval latency and recall
  • Prompt / model version
  • Source freshness and index version
  • Token usage and cost
  • Fallback / no-answer rate
  • Citation and faithfulness quality
RAG trace
trace = {
  "index_version": "hr_2026_09",
  "retriever": "hybrid_v3",
  "prompt_version": "rag_answer_v5",
  "model_version": MODEL_VERSION,
  "source_ids": [h.id for h in hits],
  "latency_ms": latency_ms
}
✅

The next training, #10 Evals, goes deeper into systematic measurement.

Practice — 3 cases

Practice 22 / Práctica 22
Which set of fields best supports RAG traceability?
Practice 23 / Práctica 23
Why version the retrieval and prompt components separately?
Practice 24 / Práctica 24
What is the strongest production mindset for RAG?
5-Question Knowledge Check

Can you design a grounded RAG system?

Open each item after answering it in your own words.

1. What are the two core stages of RAG?

Retrieval selects relevant external evidence; generation uses selected evidence as context to produce an answer.

2. Why is context assembly important?

Retrieved chunks must be selected, deduplicated, labeled and formatted so useful evidence fits the context and remains traceable.

3. What should a grounded system do when evidence is insufficient?

Follow an explicit fallback policy, state the limitation, and avoid presenting unsupported claims as established facts.

4. Why evaluate retrieval and generation separately?

Because a RAG failure may come from different layers.

5. What makes RAG production-ready?

Versioned components, evaluation, access control, observability, freshness monitoring and safe fallback behavior.

RAG Blueprint

A reusable production pattern

LayerPurpose
SourcesMaintain approved, current knowledge
IngestParse, chunk, label and version source content
IndexCreate embeddings and searchable vector / hybrid indexes
RetrieveFind eligible evidence for the user's request
AssembleSelect, deduplicate, label and format context
GenerateAnswer from evidence with clear fallback rules
EvaluateMeasure retrieval, faithfulness, relevance, citations and system behavior
ObserveTrack versions, source IDs, latency, cost, freshness and failures

Certificate of Participation

Complete at least 12 of the 24 practice cases (50%) and enter your name.

0 / 24 • 0%
Production note

RAG behavior depends on source quality, retrieval design, model behavior and workflow constraints.