Data Scientist → AI Engineer • Training 07
Article-Training • AI Engineering Foundations

Embeddings

Turning Meaning into Numbers

Learn how AI systems convert meaning into vectors, compare semantic similarity, prepare chunks, build a simple search pipeline, and evaluate retrieval before moving into vector databases and RAG.

Text → Embedding Model → Vector → Similarity → Retrieval
📝 TEXT
→
🧠 EMBED
→
[0.12, …]
🔎 QUERY
→
📐 SIMILARITY
→
🎯 TOP-K
8learning modules
24interactive practices
5rapid review questions
50%certificate threshold
Learning target
Understand embeddings well enough to design, test and operate the vector representation layer that powers semantic search, vector databases and RAG.
Practice progress0 / 24

Complete 12 of 24 practices (50%) and enter your name to unlock the Certificate of Participation.

MODULE 01
🧠

Embeddings: Meaning as Coordinates

An embedding turns text, images or other data into a numeric vector. The numbers are not a human-readable definition; together they place the item in a learned space where related meaning tends to appear nearby.

👁️
See it this way

A map uses latitude and longitude to place cities. An embedding uses many learned coordinates to place meanings.

Core ideas

  • Input becomes a fixed-length vector
  • Nearby vectors often represent related meaning
  • The coordinates are learned by a model
  • Embeddings enable comparison by geometry
Conceptual example
text = "reset my password"
embedding = [0.12, -0.44, 0.81, ..., 0.07]

# Meaning → numeric position
✅

Do not interpret one coordinate by itself. Meaning comes from the vector as a whole.

Practice — 3 cases

Practice 1 / Práctica 1
What is an embedding?
Practice 2 / Práctica 2
Why are embeddings useful for semantic search?
Practice 3 / Práctica 3
Which statement is safest about one individual embedding dimension?
MODULE 02
📐

Dimensions and the Geometry of Meaning

Embedding models output vectors with a defined dimensionality. More dimensions do not automatically mean better results; quality depends on the model, the task, the data and how retrieval is evaluated.

🗺️
A mental picture

We can draw 2D or 3D examples, but real embeddings may use hundreds or thousands of dimensions. The geometry still gives us a way to compare positions.

What matters

  • Dimensionality is part of the embedding model contract
  • All vectors compared together must be compatible
  • Distance and direction capture relationships
  • Changing embedding models usually requires re-embedding stored content
📄 policy
📄 procedure
📄 manual
⚽ soccer
🏀 basketball
🎾 tennis
[0.31, −0.08, 0.72, …, 0.15]
✅

A vector's dimensionality must match the index or storage system expecting it.

Practice — 3 cases

Practice 4 / Práctica 4
You switch to a different embedding model with a different vector size. What should you expect?
Practice 5 / Práctica 5
Does a larger vector automatically guarantee better retrieval?
Practice 6 / Práctica 6
Why do we often use 2D diagrams for embeddings?
MODULE 03
📏

Similarity: How Vectors Become Search

Once items are vectors, we can rank them by similarity or distance. Cosine similarity is common for semantic comparison; dot product and Euclidean distance are also used depending on model and index design.

🧲
See it this way

The query becomes a vector too. Search asks: which stored vectors are closest or most aligned with this query vector?

Cosine similarity

Cosine similarity focuses on vector direction.

NumPy
import numpy as np

def cosine(a, b):
    a, b = np.array(a), np.array(b)
    return np.dot(a, b) / (
        np.linalg.norm(a) * np.linalg.norm(b)
    )
Cosine
Direction / angle
Dot Product
Alignment + magnitude behavior
Euclidean
Straight-line distance

Practice — 3 cases

Practice 7 / Práctica 7
What happens first in semantic vector search?
Practice 8 / Práctica 8
Which metric compares vector direction and is common in semantic similarity?
Practice 9 / Práctica 9
Should you copy a similarity threshold from another project without testing?
MODULE 04
✂️

What You Embed Matters: Chunking and Context

Long documents are usually broken into chunks before embedding. Chunk size, overlap, headings and metadata affect whether the right evidence can be found later.

🧩
The retrieval unit

Search retrieves the units you indexed.

Chunking questions

  • Does each chunk preserve enough context?
  • Are headings or section names retained?
  • Is overlap useful or just duplicate noise?
  • What metadata will help filtering later?
Simple chunk record
chunk = {
  "text": section_text,
  "source": "employee_policy.pdf",
  "section": "Leave",
  "effective_date": "2026-01-01"
}
✅

Embedding quality and chunk quality are different problems.

Practice — 3 cases

Practice 10 / Práctica 10
A 100-page manual is stored as one giant embedding. What is a likely problem?
Practice 11 / Práctica 11
Why retain metadata with chunks?
Practice 12 / Práctica 12
What is the best chunk size for every application?
MODULE 05
🐍

Generating Embeddings in an Application

The application sends content to an embedding model and receives a vector.

🔁
Same transformation, two moments

At indexing time you embed the corpus. At search time you embed the user's query.

Application flow

Text→Embed→Vector→Store

Query→Embed→Compare→Top-K
Provider-neutral pseudocode
def embed(text):
    result = embedding_client.create(
        model=EMBEDDING_MODEL,
        input=text
    )
    return result.vector

query_vector = embed(user_query)
✅

Pin or record the embedding model/version used to create an index.

Practice — 3 cases

Practice 13 / Práctica 13
What should be stored with a document vector for traceability?
Practice 14 / Práctica 14
At search time, how should the user query be prepared?
Practice 15 / Práctica 15
Why record the embedding model used for an index?
MODULE 06
🔎

Build a Minimal Semantic Search Pipeline

A minimal semantic search system embeds records, stores vectors, embeds the query, computes similarity, and returns the highest-ranked matches.

🎯
The bridge to #08

Embeddings give us the vectors. A vector database gives us an efficient place to index and search those vectors with metadata.

Retrieval loop

  1. Embed corpus
  2. Embed query
  3. Score candidates
  4. Sort by similarity
  5. Return top-k with source metadata
Small in-memory example
query_vec = embed(query)

ranked = sorted(
    docs,
    key=lambda d: cosine(query_vec, d["vector"]),
    reverse=True
)

top_k = ranked[:3]
"reset account access"
"change password"
"street light repair"

Practice — 3 cases

Practice 16 / Práctica 16
What does top-k mean in a retrieval step?
Practice 17 / Práctica 17
Why return source metadata with retrieved chunks?
Practice 18 / Práctica 18
What is the next infrastructure step once the corpus becomes large?
MODULE 07
🧪

Evaluate Retrieval Quality, Not Just Vector Math

A technically valid similarity score does not prove that retrieval is useful.

📊
Operational question

Do not ask only whether similarity is high. Ask whether the correct evidence appeared.

Recall@K
Did relevant evidence appear in the retrieved set?
Precision
How much of what we returned was actually useful?
Latency / Cost
Can the retrieval design operate at required scale?

Common failure modes

  • Wrong chunk boundaries
  • Missing or stale documents
  • Embedding model mismatch
  • Poor metadata filters
  • Top-k too small or too large
Evaluation sketch
for case in retrieval_eval:
    results = search(case.query, k=5)
    hit = case.expected_id in [
        r["id"] for r in results
    ]
    print(case.query, hit)

Practice — 3 cases

Practice 19 / Práctica 19
Which test best evaluates a semantic search system?
Practice 20 / Práctica 20
A relevant document never appears in top-5 results. What should you investigate?
Practice 21 / Práctica 21
Why can top-k be too large?
MODULE 08
🏭

Production Embedding Patterns

Production embedding systems need versioning, observability, re-index plans, privacy controls and clear ownership of source data.

🏗️
Operate the index

Know which model created each index, when the corpus was refreshed, what sources are included, and how you will rebuild safely when models or data change.

Production checklist

  • Record embedding model/version and dimensions
  • Keep source IDs and metadata
  • Define incremental update and full rebuild paths
  • Monitor index freshness, latency and retrieval quality
  • Apply privacy and access controls to source content and retrieval
Index manifest
INDEX_MANIFEST = {
  "embedding_model": "model_v1",
  "dimensions": 1536,
  "chunker": "policy_chunks_v2",
  "last_refresh": "2026-09-23",
  "source_scope": "approved_docs"
}
✅

Treat an embedding index as a versioned data product.

Practice — 3 cases

Practice 22 / Práctica 22
What belongs in an embedding index manifest?
Practice 23 / Práctica 23
A policy document changes. What should a maintained embedding system do?
Practice 24 / Práctica 24
Which production principle is strongest?
5-Question Knowledge Check

Can you explain how meaning becomes searchable?

Open each item after answering it in your own words.

1. What does an embedding represent?

A learned numeric representation that places an item in a vector space where related items can be compared geometrically.

2. Why must query and corpus vectors be compatible?

Similarity is meaningful only when vectors belong to a compatible embedding space and dimensional setup.

3. Why does chunking affect retrieval quality?

Chunks are the retrieval units.

4. What does top-k retrieval do?

It returns the k highest-ranked matches.

5. How do you know whether embeddings are working well?

Evaluate representative user queries and verify that relevant evidence appears at useful ranks.

Embedding Blueprint

The reusable pipeline

StagePurpose
PrepareClean content, preserve useful structure and attach metadata
ChunkCreate retrieval units that preserve enough meaning
EmbedTransform each chunk into a compatible numeric vector
IndexStore vectors, source references and metadata for search
QueryEmbed the user's request with a compatible setup
RetrieveRank candidates by similarity and apply useful filters
EvaluateMeasure whether relevant evidence appears for real questions

Certificate of Participation

Complete at least 12 of the 24 practice cases (50%) and enter your name.

0 / 24 • 0%
Production note

Embedding behavior varies by model, data and task. Validate retrieval on your own corpus, record the embedding model/version and dimensions, and plan for re-indexing when the representation layer changes.