Build the conceptual foundation for working with large language models without magic thinking: understand tokenization, next-token generation, context, embeddings, sampling, limitations and the product patterns built on top.
Complete 12 of 24 practices (50%) and enter your name to unlock the Certificate of Participation.
A large language model is a neural network trained on large text/code corpora to model patterns in sequences. During inference it generates one token at a time based on the context it receives.
Treat the model as a powerful generator and reasoner over provided context—not as an infallible database.
prompt = "Explain churn in plain English"
# LLM estimates a probability distribution
# for the next token, then continues.Treat the model as a powerful generator and reasoner over provided context—not as an infallible database.
Models do not read words directly. Tokenizers convert text into token IDs, often using subword pieces. Token count affects context capacity, latency and cost.
Estimate and measure tokens when designing long prompts, document ingestion and cost-sensitive workflows.
text = "Data science meets AI engineering"
# tokenizer(text) -> token ids
# [ ... model-specific integers ... ]Estimate and measure tokens when designing long prompts, document ingestion and cost-sensitive workflows.
Most chat LLMs generate sequentially. At each step the model predicts probabilities for the next token, a decoding strategy selects one, and the new token becomes part of the context for the next step.
Fluent output can still be wrong because generation optimizes likely continuation, not guaranteed factual truth.
context = prompt
while not done:
probs = model.next_token(context)
token = sample(probs)
context += tokenFluent output can still be wrong because generation optimizes likely continuation, not guaranteed factual truth.
The context window is the amount of tokenized information available to the model during a request. Attention lets the model weigh relationships among tokens, but more context is not automatically better context.
Curate context. Irrelevant or contradictory information can reduce quality even when the model supports a large window.
context = [
system_instructions,
conversation_history,
retrieved_evidence,
current_question
]Curate context. Irrelevant or contradictory information can reduce quality even when the model supports a large window.
Embeddings map text (or other content) into numeric vectors where semantically related items tend to be closer. They power similarity search, clustering and retrieval pipelines.
Embeddings help find relevant context; the LLM then uses that context to generate or reason.
query_vec = embed("late invoice risk")
doc_vecs = embed(documents)
# similarity(query_vec, doc_vecs)Embeddings help find relevant context; the LLM then uses that context to generate or reason.
Decoding parameters shape how tokens are selected. Lower randomness favors more predictable continuations; higher randomness increases diversity but can also increase instability.
Use lower randomness for extraction and controlled workflows; allow more creativity only when the task benefits from it.
settings = {
"temperature": 0.2,
"top_p": 0.9,
"max_output_tokens": 500
}Use lower randomness for extraction and controlled workflows; allow more creativity only when the task benefits from it.
LLMs can produce plausible but unsupported statements, may not contain current/private knowledge, and inherit biases from training and product context. These are engineering constraints, not edge cases.
The correct response to model uncertainty is system design: retrieval, tools, validation, citations, human review and evals.
answer = llm(question)
# Do not trust by default.
# Ground, verify, evaluate, and constrain.The correct response to model uncertainty is system design: retrieval, tools, validation, citations, human review and evals.
A production LLM application wraps the model with prompts, context management, retrieval, tools, guardrails, evaluation, observability and business logic. The model is a component—not the whole product.
Once you understand the foundation, RAG, agents and orchestration stop looking like magic—they become engineering patterns around a probabilistic model.
user -> app logic -> LLM
-> retrieval
-> tools/APIs
-> guardrails
-> evaluated responseOnce you understand the foundation, RAG, agents and orchestration stop looking like magic—they become engineering patterns around a probabilistic model.
Open each item only after answering it in your own words.
A sequence of tokens, selected one step at a time from probability distributions conditioned on the current context.
They affect context capacity, latency and often usage cost.
The token budget of information the model can consider during a request.
Representing semantic content as vectors for similarity, clustering and retrieval.
Real applications need business logic, context, tools, security, evaluation, observability and reliability around the model.
| Need | Recommended approach |
|---|---|
| Represent text for similarity | Embeddings |
| Control available information | Context management |
| Reduce unsupported answers | Grounding + verification |
| Build production app | LLM + surrounding system |
Complete at least 12 of the 24 practice cases (50%) and enter your name.