Data Scientist → AI Engineer • Training 08
Article-Training • AI Engineering Foundations

Vector Databases

Storing and Searching Meaning at Scale

Move from individual embeddings to a production retrieval layer. Learn how vector databases store vectors with metadata, build similarity indexes, filter results, support hybrid search, manage updates, and serve the retrieval step that powers RAG.

Documents → Embeddings → Vector Index → Query → Top-K Results → RAG
📄 DOCS
→
🧠 EMBED
→
🗄️ VECTOR DB
🔎 QUERY
→
📐 SEARCH
→
🎯 TOP-K
8learning modules
24interactive practices
5rapid review questions
50%certificate threshold
Learning target
Design and operate a vector retrieval layer that stores embeddings with metadata, returns useful candidates quickly, and remains measurable, refreshable and ready for RAG.
Practice progress0 / 24

Complete 12 of 24 practices (50%) and enter your name to unlock the Certificate of Participation.

MODULE 01
🗄️

What a Vector Database Actually Adds

Embeddings give us vectors. A vector database gives us an operational system for storing those vectors, associating them with IDs and metadata, indexing them, and retrieving the nearest candidates efficiently.

👁️
See it this way

A spreadsheet can hold vectors. A vector database is built to search millions of them quickly while preserving metadata, filters and update operations.

Core pieces

  • Vector + stable record ID
  • Source text or durable source reference
  • Metadata for filters and traceability
  • Similarity index for fast nearest-neighbor search
IDpolicy_042
VECTOR[0.14, −0.08, …]
METADATAdept=HR • year=2026
SOURCELeave Policy §4
✅

A vector database is not the embedding model. It stores and searches the vectors produced by an embedding model.

Practice — 3 cases

Practice 1 / Práctica 1
What is the main role of a vector database?
Practice 2 / Práctica 2
Which record design is strongest for retrieval?
Practice 3 / Práctica 3
Which statement correctly separates embeddings from vector databases?
MODULE 02
⚡

Indexing and Approximate Nearest Neighbors

Comparing a query against every stored vector can become expensive at scale. Vector indexes use nearest-neighbor strategies—often approximate—to trade a small amount of recall for much lower latency and higher throughput.

🏙️
Think like a city map

If you need the nearest coffee shop, you do not measure the distance to every business on Earth. An index narrows the search intelligently.

Trade-offs

  • Exact search: strongest exhaustive comparison, higher cost at scale
  • Approximate search: much faster, may miss some true neighbors
  • Index parameters influence recall, latency and memory
  • Tune with realistic workload, not defaults alone
Conceptual query
results = vector_index.search(
    vector=query_vector,
    top_k=5,
    search_effort="balanced"
)
Higher search effort
Lower latency

Practice — 3 cases

Practice 4 / Práctica 4
Why use an approximate nearest-neighbor index?
Practice 5 / Práctica 5
What should drive index tuning?
Practice 6 / Práctica 6
Which statement about exact search is most accurate?
MODULE 03
🏷️

IDs, Metadata, Filters and Namespaces

Semantic similarity alone is rarely enough. Real systems also need filters such as department, date, language, security scope, document type or tenant. Metadata turns vector search into controlled retrieval.

🔐
Similarity is not permission

A document can be semantically perfect and still be the wrong result if the user is not authorized to see it or if it belongs to another tenant.

Useful metadata

  • source_id / document_id
  • department / branch / tenant
  • effective_date / version
  • language / document_type
  • security_scope / access_group
Filtered retrieval
results = db.search(
    vector=query_vector,
    top_k=8,
    filter={
        "department": "HR",
        "status": "active"
    }
)
✅

Use metadata filters to narrow eligible candidates before or during similarity search. Authorization should still be enforced by deterministic application controls.

Practice — 3 cases

Practice 7 / Práctica 7
Why attach metadata to vectors?
Practice 8 / Práctica 8
A user's query is most similar to a confidential document they cannot access. What should happen?
Practice 9 / Práctica 9
What is a namespace or tenant partition useful for?
MODULE 04
🔎

The Query Path: From Question to Top-K Results

At query time, the application embeds the user's request, applies eligibility filters, searches the vector index, ranks candidates, and returns top results with source information.

🎯
Retrieval is a pipeline

The score is only one step. A production query also needs filters, top-k policy, source hydration and often reranking before the context reaches an LLM.

Query→Embed→Filter→Vector Search→Top-K→Fetch Source

Query controls

  • top_k: how many candidates to return
  • minimum score or acceptance logic
  • metadata / access filters
  • optional reranking stage
Provider-neutral pseudocode
q = embed(user_query)

hits = vector_db.search(
    vector=q,
    top_k=10,
    filter=access_filter
)

context = hydrate_sources(hits[:5])

Practice — 3 cases

Practice 10 / Práctica 10
What does top-k control?
Practice 11 / Práctica 11
Why fetch the source text after vector search?
Practice 12 / Práctica 12
When is reranking useful?
MODULE 05
🔀

Hybrid Search: Keywords + Vectors

Vector search is strong at meaning; lexical search is strong at exact terms such as product codes, names, acronyms and legal phrases. Hybrid retrieval combines both signals and often improves real-world search.

🤝
Two strengths, one ranking

A query such as 'Form HR-17 parental leave' contains both semantic intent and an exact identifier. A hybrid system can preserve both.

Search signals

SignalBest at
VectorMeaning, paraphrases, conceptual similarity
LexicalExact terms, codes, names, uncommon tokens
HybridCombining both signals into a stronger candidate ranking
Conceptual hybrid query
results = search(
    text_query=query,
    vector=query_vector,
    alpha=0.65,   # vector weight
    top_k=10
)
✅

Hybrid search is not automatically better. Measure it against a retrieval evaluation set and tune weighting for your corpus.

Practice — 3 cases

Practice 13 / Práctica 13
When can lexical search outperform pure vector search?
Practice 14 / Práctica 14
What does hybrid search combine?
Practice 15 / Práctica 15
How should hybrid weighting be chosen?
MODULE 06
🔁

Upserts, Updates, Deletes and Re-Indexing

A vector database is a living data product. Documents change, policies expire, records are deleted, and embedding models evolve. Production systems need explicit lifecycle operations.

🧹
Fresh retrieval requires fresh indexes

If a policy changed yesterday but the old vector is still active, your retrieval layer can confidently return stale evidence.

Lifecycle operations

  • Upsert: insert or replace a known record
  • Delete: remove obsolete or unauthorized records
  • Refresh: re-embed changed source content
  • Rebuild: create a new index when model/schema/index design changes materially
Incremental refresh
for doc in changed_documents:
    chunks = chunk(doc)
    vectors = embed(chunks)
    vector_db.upsert(vectors)

for doc_id in deleted_ids:
    vector_db.delete(filter={
        "source_id": doc_id
    })

Practice — 3 cases

Practice 16 / Práctica 16
A source document changes but keeps the same business identity. What is usually appropriate?
Practice 17 / Práctica 17
When is a full re-index especially likely to be needed?
Practice 18 / Práctica 18
Why must deletes propagate to the vector index?
MODULE 07
📊

Evaluate the Retrieval Layer

A vector database is valuable only if it returns useful evidence fast enough and within the correct data scope. Evaluate retrieval quality and operational behavior together.

🧪
Measure the system, not a demo

One impressive query proves little. Track retrieval metrics across common, difficult, filtered and adversarial cases.

Recall@K
Did relevant evidence appear in the candidate set?
P95 Latency
How slow are the slower real requests?
Freshness
Is the index synchronized with source changes?

Also monitor

  • Filter / authorization correctness
  • Index size and memory footprint
  • Query throughput
  • No-result / low-confidence rates
Retrieval test loop
for case in eval_set:
    hits = search(case.query, k=5)
    recall_hit = case.expected_id in [
        h["source_id"] for h in hits
    ]
    record(case.id, recall_hit)

Practice — 3 cases

Practice 19 / Práctica 19
Which metric directly asks whether relevant evidence appeared in the first k results?
Practice 20 / Práctica 20
Why monitor index freshness?
Practice 21 / Práctica 21
What does P95 query latency help reveal?
MODULE 08
🏭

Production Architecture — The Retrieval Layer Before RAG

By itself, a vector database returns candidates. In a RAG system, those candidates become grounded context for an LLM. A reliable architecture separates ingestion, indexing, retrieval, access control, evaluation and generation.

🌉
The bridge to #09

Embeddings represent meaning. Vector databases retrieve relevant meaning. RAG takes that retrieved evidence and places it into the model's context so the LLM can answer from external knowledge.

Sources→Chunk→Embed→Vector DB→Retrieve→Context→LLM

Production responsibilities

  • Ingestion and document freshness
  • Embedding / index version compatibility
  • Authorization-aware retrieval
  • Retrieval evaluation and observability
  • Graceful behavior when no strong evidence is found
Architecture sketch
query_vec = embed(question)
hits = vector_db.search(
    query_vec,
    filter=user_scope,
    top_k=8
)

context = rerank_and_format(hits)
answer = llm(question, context)
✅

The vector database does not guarantee a correct answer. It improves the evidence selection step. RAG quality still depends on retrieval, prompting, model behavior, evaluation and guardrails.

Practice — 3 cases

Practice 22 / Práctica 22
What does a vector database contribute to a RAG pipeline?
Practice 23 / Práctica 23
What should happen when retrieval finds no strong evidence?
Practice 24 / Práctica 24
Which architecture is strongest for production?
5-Question Knowledge Check

Can you design a production vector retrieval layer?

Open each item after answering it in your own words. The 24 interactive practices above drive certificate progress.

1. What does a vector database add beyond embeddings?

Operational storage, IDs, metadata, similarity indexes, filters, updates and fast retrieval over many vectors.

2. Why use approximate nearest-neighbor indexing?

To lower latency and search cost at scale while accepting and measuring a possible recall trade-off.

3. Why is metadata essential in production retrieval?

It enables source traceability, freshness checks, filtering, tenant separation and authorization-aware candidate selection.

4. When is hybrid search useful?

When semantic meaning and exact lexical signals such as codes, names or rare terms both matter.

5. How does a vector database connect to RAG?

It retrieves relevant external evidence; RAG then injects selected evidence into the LLM context for grounded generation.

Vector Database Blueprint

A reusable production pattern

LayerPurpose
IngestCollect approved source documents and changes
Chunk + EmbedCreate retrieval units and compatible vectors
StorePersist stable IDs, vectors, source references and metadata
IndexBuild a similarity structure tuned for recall, latency and memory
FilterRestrict candidates by tenant, access scope, date or document type
RetrieveReturn top-k candidates and source evidence
EvaluateMeasure recall, latency, freshness and filter correctness
RefreshUpsert, delete or rebuild as source data and embedding models change

Certificate of Participation

Complete at least 12 of the 24 practice cases (50%) and enter your name.

0 / 24 • 0%
Production note

Vector database behavior depends on index type, corpus size, filters, embedding model and workload. Validate recall, latency, access control and freshness on the actual system you plan to operate. Product-specific APIs and tuning parameters vary by provider.