Move from individual embeddings to a production retrieval layer. Learn how vector databases store vectors with metadata, build similarity indexes, filter results, support hybrid search, manage updates, and serve the retrieval step that powers RAG.
Complete 12 of 24 practices (50%) and enter your name to unlock the Certificate of Participation.
Embeddings give us vectors. A vector database gives us an operational system for storing those vectors, associating them with IDs and metadata, indexing them, and retrieving the nearest candidates efficiently.
A spreadsheet can hold vectors. A vector database is built to search millions of them quickly while preserving metadata, filters and update operations.
A vector database is not the embedding model. It stores and searches the vectors produced by an embedding model.
Comparing a query against every stored vector can become expensive at scale. Vector indexes use nearest-neighbor strategies—often approximate—to trade a small amount of recall for much lower latency and higher throughput.
If you need the nearest coffee shop, you do not measure the distance to every business on Earth. An index narrows the search intelligently.
results = vector_index.search(
vector=query_vector,
top_k=5,
search_effort="balanced"
)Semantic similarity alone is rarely enough. Real systems also need filters such as department, date, language, security scope, document type or tenant. Metadata turns vector search into controlled retrieval.
A document can be semantically perfect and still be the wrong result if the user is not authorized to see it or if it belongs to another tenant.
results = db.search(
vector=query_vector,
top_k=8,
filter={
"department": "HR",
"status": "active"
}
)Use metadata filters to narrow eligible candidates before or during similarity search. Authorization should still be enforced by deterministic application controls.
At query time, the application embeds the user's request, applies eligibility filters, searches the vector index, ranks candidates, and returns top results with source information.
The score is only one step. A production query also needs filters, top-k policy, source hydration and often reranking before the context reaches an LLM.
q = embed(user_query)
hits = vector_db.search(
vector=q,
top_k=10,
filter=access_filter
)
context = hydrate_sources(hits[:5])Vector search is strong at meaning; lexical search is strong at exact terms such as product codes, names, acronyms and legal phrases. Hybrid retrieval combines both signals and often improves real-world search.
A query such as 'Form HR-17 parental leave' contains both semantic intent and an exact identifier. A hybrid system can preserve both.
| Signal | Best at |
|---|---|
| Vector | Meaning, paraphrases, conceptual similarity |
| Lexical | Exact terms, codes, names, uncommon tokens |
| Hybrid | Combining both signals into a stronger candidate ranking |
results = search(
text_query=query,
vector=query_vector,
alpha=0.65, # vector weight
top_k=10
)Hybrid search is not automatically better. Measure it against a retrieval evaluation set and tune weighting for your corpus.
A vector database is a living data product. Documents change, policies expire, records are deleted, and embedding models evolve. Production systems need explicit lifecycle operations.
If a policy changed yesterday but the old vector is still active, your retrieval layer can confidently return stale evidence.
for doc in changed_documents:
chunks = chunk(doc)
vectors = embed(chunks)
vector_db.upsert(vectors)
for doc_id in deleted_ids:
vector_db.delete(filter={
"source_id": doc_id
})A vector database is valuable only if it returns useful evidence fast enough and within the correct data scope. Evaluate retrieval quality and operational behavior together.
One impressive query proves little. Track retrieval metrics across common, difficult, filtered and adversarial cases.
for case in eval_set:
hits = search(case.query, k=5)
recall_hit = case.expected_id in [
h["source_id"] for h in hits
]
record(case.id, recall_hit)By itself, a vector database returns candidates. In a RAG system, those candidates become grounded context for an LLM. A reliable architecture separates ingestion, indexing, retrieval, access control, evaluation and generation.
Embeddings represent meaning. Vector databases retrieve relevant meaning. RAG takes that retrieved evidence and places it into the model's context so the LLM can answer from external knowledge.
query_vec = embed(question)
hits = vector_db.search(
query_vec,
filter=user_scope,
top_k=8
)
context = rerank_and_format(hits)
answer = llm(question, context)The vector database does not guarantee a correct answer. It improves the evidence selection step. RAG quality still depends on retrieval, prompting, model behavior, evaluation and guardrails.
Open each item after answering it in your own words. The 24 interactive practices above drive certificate progress.
Operational storage, IDs, metadata, similarity indexes, filters, updates and fast retrieval over many vectors.
To lower latency and search cost at scale while accepting and measuring a possible recall trade-off.
It enables source traceability, freshness checks, filtering, tenant separation and authorization-aware candidate selection.
When semantic meaning and exact lexical signals such as codes, names or rare terms both matter.
It retrieves relevant external evidence; RAG then injects selected evidence into the LLM context for grounded generation.
| Layer | Purpose |
|---|---|
| Ingest | Collect approved source documents and changes |
| Chunk + Embed | Create retrieval units and compatible vectors |
| Store | Persist stable IDs, vectors, source references and metadata |
| Index | Build a similarity structure tuned for recall, latency and memory |
| Filter | Restrict candidates by tenant, access scope, date or document type |
| Retrieve | Return top-k candidates and source evidence |
| Evaluate | Measure recall, latency, freshness and filter correctness |
| Refresh | Upsert, delete or rebuild as source data and embedding models change |
Complete at least 12 of the 24 practice cases (50%) and enter your name.
Vector database behavior depends on index type, corpus size, filters, embedding model and workload. Validate recall, latency, access control and freshness on the actual system you plan to operate. Product-specific APIs and tuning parameters vary by provider.