Learn how AI systems convert meaning into vectors, compare semantic similarity, prepare chunks, build a simple search pipeline, and evaluate retrieval before moving into vector databases and RAG.
Complete 12 of 24 practices (50%) and enter your name to unlock the Certificate of Participation.
An embedding turns text, images or other data into a numeric vector. The numbers are not a human-readable definition; together they place the item in a learned space where related meaning tends to appear nearby.
A map uses latitude and longitude to place cities. An embedding uses many learned coordinates to place meanings.
text = "reset my password"
embedding = [0.12, -0.44, 0.81, ..., 0.07]
# Meaning → numeric positionDo not interpret one coordinate by itself. Meaning comes from the vector as a whole.
Embedding models output vectors with a defined dimensionality. More dimensions do not automatically mean better results; quality depends on the model, the task, the data and how retrieval is evaluated.
We can draw 2D or 3D examples, but real embeddings may use hundreds or thousands of dimensions. The geometry still gives us a way to compare positions.
A vector's dimensionality must match the index or storage system expecting it.
Once items are vectors, we can rank them by similarity or distance. Cosine similarity is common for semantic comparison; dot product and Euclidean distance are also used depending on model and index design.
The query becomes a vector too. Search asks: which stored vectors are closest or most aligned with this query vector?
Cosine similarity focuses on vector direction.
import numpy as np
def cosine(a, b):
a, b = np.array(a), np.array(b)
return np.dot(a, b) / (
np.linalg.norm(a) * np.linalg.norm(b)
)Long documents are usually broken into chunks before embedding. Chunk size, overlap, headings and metadata affect whether the right evidence can be found later.
Search retrieves the units you indexed.
chunk = {
"text": section_text,
"source": "employee_policy.pdf",
"section": "Leave",
"effective_date": "2026-01-01"
}Embedding quality and chunk quality are different problems.
The application sends content to an embedding model and receives a vector.
At indexing time you embed the corpus. At search time you embed the user's query.
def embed(text):
result = embedding_client.create(
model=EMBEDDING_MODEL,
input=text
)
return result.vector
query_vector = embed(user_query)Pin or record the embedding model/version used to create an index.
A minimal semantic search system embeds records, stores vectors, embeds the query, computes similarity, and returns the highest-ranked matches.
Embeddings give us the vectors. A vector database gives us an efficient place to index and search those vectors with metadata.
query_vec = embed(query)
ranked = sorted(
docs,
key=lambda d: cosine(query_vec, d["vector"]),
reverse=True
)
top_k = ranked[:3]A technically valid similarity score does not prove that retrieval is useful.
Do not ask only whether similarity is high. Ask whether the correct evidence appeared.
for case in retrieval_eval:
results = search(case.query, k=5)
hit = case.expected_id in [
r["id"] for r in results
]
print(case.query, hit)Production embedding systems need versioning, observability, re-index plans, privacy controls and clear ownership of source data.
Know which model created each index, when the corpus was refreshed, what sources are included, and how you will rebuild safely when models or data change.
INDEX_MANIFEST = {
"embedding_model": "model_v1",
"dimensions": 1536,
"chunker": "policy_chunks_v2",
"last_refresh": "2026-09-23",
"source_scope": "approved_docs"
}Treat an embedding index as a versioned data product.
Open each item after answering it in your own words.
A learned numeric representation that places an item in a vector space where related items can be compared geometrically.
Similarity is meaningful only when vectors belong to a compatible embedding space and dimensional setup.
Chunks are the retrieval units.
It returns the k highest-ranked matches.
Evaluate representative user queries and verify that relevant evidence appears at useful ranks.
| Stage | Purpose |
|---|---|
| Prepare | Clean content, preserve useful structure and attach metadata |
| Chunk | Create retrieval units that preserve enough meaning |
| Embed | Transform each chunk into a compatible numeric vector |
| Index | Store vectors, source references and metadata for search |
| Query | Embed the user's request with a compatible setup |
| Retrieve | Rank candidates by similarity and apply useful filters |
| Evaluate | Measure whether relevant evidence appears for real questions |
Complete at least 12 of the 24 practice cases (50%) and enter your name.
Embedding behavior varies by model, data and task. Validate retrieval on your own corpus, record the embedding model/version and dimensions, and plan for re-indexing when the representation layer changes.