Data Scientist → AI Engineer • Training 05
Article-Training • AI Engineering Foundations

LLM Foundations

Tokens, Context, Embeddings, Inference and the Boundaries of Generation

Build the conceptual foundation for working with large language models without magic thinking: understand tokenization, next-token generation, context, embeddings, sampling, limitations and the product patterns built on top.

Text → Tokens → Model → Probabilities → Sampling → Generated Output
📝 TEXT
→
🔤 TOKENS
→
🧠 LLM
→
✨ OUTPUT
8learning modules
24interactive practices
5rapid review questions
50%certificate threshold
Learning target
Understand what LLMs are, how tokenization and autoregressive generation work, how context windows shape behavior, what embeddings represent, how sampling changes output and why grounding/evaluation are necessary.
Practice progress0 / 24

Complete 12 of 24 practices (50%) and enter your name to unlock the Certificate of Participation.

MODULE 01
🧭

What an LLM Is — and Is Not

A large language model is a neural network trained on large text/code corpora to model patterns in sequences. During inference it generates one token at a time based on the context it receives.

👁️
See it this way

Treat the model as a powerful generator and reasoner over provided context—not as an infallible database.

Core ideas

  • Neural sequence model
  • Learns statistical language patterns
  • Generates token by token
  • Does not automatically know your private/current data
Try this
prompt = "Explain churn in plain English"
# LLM estimates a probability distribution
# for the next token, then continues.
✅

Treat the model as a powerful generator and reasoner over provided context—not as an infallible database.

Practice — 3 cases

Practice 1 / Práctica 1
Which statement best reflects this module?
Practice 2 / Práctica 2
Which action is the best engineering choice?
Practice 3 / Práctica 3
Which option would you avoid in a production system?
MODULE 02
🧱

Tokenization: How Text Becomes Model Input

Models do not read words directly. Tokenizers convert text into token IDs, often using subword pieces. Token count affects context capacity, latency and cost.

👁️
See it this way

Estimate and measure tokens when designing long prompts, document ingestion and cost-sensitive workflows.

Core ideas

  • Text is segmented into tokens
  • Tokens map to integer IDs
  • Different models use different tokenizers
  • Token count is not equal to word count
Try this
text = "Data science meets AI engineering"
# tokenizer(text) -> token ids
# [ ... model-specific integers ... ]
✅

Estimate and measure tokens when designing long prompts, document ingestion and cost-sensitive workflows.

Practice — 3 cases

Practice 4 / Práctica 4
Which statement best reflects this module?
Practice 5 / Práctica 5
Which action is the best engineering choice?
Practice 6 / Práctica 6
Which option would you avoid in a production system?
MODULE 03
🛠️

Autoregressive Generation

Most chat LLMs generate sequentially. At each step the model predicts probabilities for the next token, a decoding strategy selects one, and the new token becomes part of the context for the next step.

👁️
See it this way

Fluent output can still be wrong because generation optimizes likely continuation, not guaranteed factual truth.

Core ideas

  • Predict distribution
  • Select token
  • Append to context
  • Repeat until stop condition
Try this
context = prompt
while not done:
    probs = model.next_token(context)
    token = sample(probs)
    context += token
✅

Fluent output can still be wrong because generation optimizes likely continuation, not guaranteed factual truth.

Practice — 3 cases

Practice 7 / Práctica 7
Which statement best reflects this module?
Practice 8 / Práctica 8
Which action is the best engineering choice?
Practice 9 / Práctica 9
Which option would you avoid in a production system?
MODULE 04
🧪

Context Windows & Attention

The context window is the amount of tokenized information available to the model during a request. Attention lets the model weigh relationships among tokens, but more context is not automatically better context.

👁️
See it this way

Curate context. Irrelevant or contradictory information can reduce quality even when the model supports a large window.

Core ideas

  • System instructions
  • Conversation history
  • Retrieved documents
  • User request all compete for context space
Try this
context = [
  system_instructions,
  conversation_history,
  retrieved_evidence,
  current_question
]
✅

Curate context. Irrelevant or contradictory information can reduce quality even when the model supports a large window.

Practice — 3 cases

Practice 10 / Práctica 10
Which statement best reflects this module?
Practice 11 / Práctica 11
Which action is the best engineering choice?
Practice 12 / Práctica 12
Which option would you avoid in a production system?
MODULE 05
⚙️

Embeddings: Meaning as Vectors

Embeddings map text (or other content) into numeric vectors where semantically related items tend to be closer. They power similarity search, clustering and retrieval pipelines.

👁️
See it this way

Embeddings help find relevant context; the LLM then uses that context to generate or reason.

Core ideas

  • Encode content to vectors
  • Compare similarity
  • Retrieve related items
  • Not the same as generative output
Try this
query_vec = embed("late invoice risk")
doc_vecs = embed(documents)
# similarity(query_vec, doc_vecs)
✅

Embeddings help find relevant context; the LLM then uses that context to generate or reason.

Practice — 3 cases

Practice 13 / Práctica 13
Which statement best reflects this module?
Practice 14 / Práctica 14
Which action is the best engineering choice?
Practice 15 / Práctica 15
Which option would you avoid in a production system?
MODULE 06
📡

Sampling: Temperature & Determinism

Decoding parameters shape how tokens are selected. Lower randomness favors more predictable continuations; higher randomness increases diversity but can also increase instability.

👁️
See it this way

Use lower randomness for extraction and controlled workflows; allow more creativity only when the task benefits from it.

Core ideas

  • Temperature changes distribution sharpness
  • Top-p limits candidate mass
  • Seeds may improve reproducibility when supported
  • Determinism depends on provider/runtime
Try this
settings = {
  "temperature": 0.2,
  "top_p": 0.9,
  "max_output_tokens": 500
}
✅

Use lower randomness for extraction and controlled workflows; allow more creativity only when the task benefits from it.

Practice — 3 cases

Practice 16 / Práctica 16
Which statement best reflects this module?
Practice 17 / Práctica 17
Which action is the best engineering choice?
Practice 18 / Práctica 18
Which option would you avoid in a production system?
MODULE 07
🛡️

Limitations: Hallucination, Freshness & Bias

LLMs can produce plausible but unsupported statements, may not contain current/private knowledge, and inherit biases from training and product context. These are engineering constraints, not edge cases.

👁️
See it this way

The correct response to model uncertainty is system design: retrieval, tools, validation, citations, human review and evals.

Core ideas

  • Hallucination = unsupported generation
  • Training knowledge has limits
  • Private data requires explicit access
  • Bias and safety need evaluation
Try this
answer = llm(question)
# Do not trust by default.
# Ground, verify, evaluate, and constrain.
✅

The correct response to model uncertainty is system design: retrieval, tools, validation, citations, human review and evals.

Practice — 3 cases

Practice 19 / Práctica 19
Which statement best reflects this module?
Practice 20 / Práctica 20
Which action is the best engineering choice?
Practice 21 / Práctica 21
Which option would you avoid in a production system?
MODULE 08
🚀

From Foundation Model to AI Product

A production LLM application wraps the model with prompts, context management, retrieval, tools, guardrails, evaluation, observability and business logic. The model is a component—not the whole product.

👁️
See it this way

Once you understand the foundation, RAG, agents and orchestration stop looking like magic—they become engineering patterns around a probabilistic model.

Core ideas

  • Prompt/instructions
  • Context & retrieval
  • Tools and structured outputs
  • Guardrails, evals and monitoring
Try this
user -> app logic -> LLM
                 -> retrieval
                 -> tools/APIs
                 -> guardrails
                 -> evaluated response
✅

Once you understand the foundation, RAG, agents and orchestration stop looking like magic—they become engineering patterns around a probabilistic model.

Practice — 3 cases

Practice 22 / Práctica 22
Which statement best reflects this module?
Practice 23 / Práctica 23
Which action is the best engineering choice?
Practice 24 / Práctica 24
Which option would you avoid in a production system?
5-Question Knowledge Check

Can you explain the core ideas clearly?

Open each item only after answering it in your own words.

1. What does an LLM generate during inference?

A sequence of tokens, selected one step at a time from probability distributions conditioned on the current context.

2. Why do tokens matter operationally?

They affect context capacity, latency and often usage cost.

3. What is a context window?

The token budget of information the model can consider during a request.

4. What are embeddings mainly used for?

Representing semantic content as vectors for similarity, clustering and retrieval.

5. Why is an LLM not the whole AI product?

Real applications need business logic, context, tools, security, evaluation, observability and reliability around the model.

Decision Guide

What should you reach for?

NeedRecommended approach
Represent text for similarityEmbeddings
Control available informationContext management
Reduce unsupported answersGrounding + verification
Build production appLLM + surrounding system

Certificate of Participation

Complete at least 12 of the 24 practice cases (50%) and enter your name.

0 / 24 • 0%