Data Engineering for Humans • Training 10
Article-Training • End-to-End Capstone

End-to-End Data Engineering Capstone

Build a City Service Requests Pipeline from Source to Trusted Decisions

Topics 1–9 taught the pieces. Topic 10 connects them into one complete data product.

Source → Ingestion → Storage → Transformation → Orchestration → Quality → Governance → Platform → Analytics / AI
Service Requests
→
Trusted Pipeline
BI
|
API
|
AI
8capstone build modules
24interactive practices
5rapid review questions
50%certificate threshold
Learning target
Design and connect the complete data engineering lifecycle.
Practice progress0 / 24
MODULE 01
🧭

Capstone Mission & Architecture

This capstone connects the entire Data Engineering for Humans series into one practical system: a City Service Requests pipeline that receives requests, stores them, transforms them, monitors them and serves trusted analytics.

✨
Key idea

One business case • One architecture • Nine disciplines working together

Core ideas

  • Start from the business question and service outcome
  • Map every stage from source to consumer
  • Design trust, security and observability from the beginning
Capstone flow
Source → Ingestion → Storage → Transform → Orchestrate → Validate → Govern → Serve
✅

The capstone is not a summary. It is the integrated application of Topics 1–9.

Practice — 3 cases

Practice 1 / Práctica 1
What makes this a capstone instead of a summary?
Practice 2 / Práctica 2
What should define the architecture first?
Practice 3 / Práctica 3
Which view best represents an end-to-end data system?
MODULE 02
📥

Source & Ingestion

Our source is a stream of service requests with request ID, category, location, timestamps, status and channel. The ingestion layer must capture new data reliably without duplicating events.

✨
Key idea

API/File → Landing Zone → Idempotent Load → Raw History

Core ideas

  • Preserve source identifiers and ingestion timestamps
  • Use idempotent loads to protect reruns
  • Quarantine malformed records instead of silently dropping them
Python ingestion sketch
for row in incoming_requests:
    if not already_loaded(row["request_id"], row["updated_at"]):
        write_raw(row, ingested_at=utcnow())
✅

Raw ingestion should preserve evidence. Cleaning belongs downstream.

Practice — 3 cases

Practice 4 / Práctica 4
Why keep a landing/raw layer?
Practice 5 / Práctica 5
What protects a rerun from creating duplicate service requests?
Practice 6 / Práctica 6
A malformed record arrives. What is the safest first action?
MODULE 03
🗄️

Storage & Data Model

The pipeline separates raw history from curated analytics. A dimensional model turns operational requests into a clean fact table with dimensions such as service, date, location, channel and status.

✨
Key idea

Raw History → Curated Tables → Fact Requests + Dimensions

Core ideas

  • Keep immutable raw history when possible
  • Use stable business keys and surrogate keys appropriately
  • Model analytics for clear measures and reusable dimensions
SQL model sketch
FactServiceRequest(
  RequestKey, ServiceKey, DateKey, LocationKey,
  OpenedAt, ClosedAt, ResolutionMinutes, Status
)
✅

Operational storage and analytical storage solve different problems.

Practice — 3 cases

Practice 7 / Práctica 7
Where should original source payloads normally be preserved?
Practice 8 / Práctica 8
What is the central analytical event in this case?
Practice 9 / Práctica 9
Why create dimensions for service, date and location?
MODULE 04
⚙️

Transformation & Business Logic

Transformation converts raw operational events into business-ready data: standardized categories, valid timestamps, derived resolution time, current status and documented rules for reopened or canceled requests.

✨
Key idea

Clean → Standardize → Derive → Join → Validate → Publish

Core ideas

  • Make business rules explicit and testable
  • Separate reusable transformations from presentation logic
  • Keep transformations deterministic where possible
SQL transformation example
SELECT
  RequestID,
  UPPER(TRIM(Category)) AS Category,
  DATEDIFF(minute, OpenedAt, ClosedAt) AS ResolutionMinutes
FROM RawServiceRequest;
✅

A metric is trustworthy only when its business rule is clear and repeatable.

Practice — 3 cases

Practice 10 / Práctica 10
Where should the definition of ResolutionMinutes live?
Practice 11 / Práctica 11
Why standardize service categories?
Practice 12 / Práctica 12
A transformation gives different results for the same unchanged input. What property is missing?
MODULE 05
🔄

Orchestration & Automation

The workflow coordinates ingestion, transformation, quality checks and publishing. Dependencies, retries, schedules and rerun behavior must be explicit so automation remains predictable.

✨
Key idea

Ingest → Transform → Test → Publish → Notify

Core ideas

  • Model dependencies as a DAG or equivalent workflow
  • Retry transient failures without duplicating business events
  • Do not publish downstream data before quality gates pass
Workflow sketch
ingest
  >> transform
  >> quality_checks
  >> publish_semantic_model
  >> notify_success
✅

Automation should make failure states visible, not merely make jobs run.

Practice — 3 cases

Practice 13 / Práctica 13
Which task should run before publishing the semantic model?
Practice 14 / Práctica 14
Why model dependencies explicitly?
Practice 15 / Práctica 15
A transient API timeout occurs. What is usually appropriate?
MODULE 06
🩺

Quality & Observability

A production pipeline needs evidence that data is complete, fresh and logically valid. Observability adds metrics, logs, lineage and alerts so the team can detect and diagnose failures quickly.

✨
Key idea

Tests + Freshness + Volume + Logs + Lineage + Alerts

Core ideas

  • Test critical keys, nulls, domains and business invariants
  • Monitor freshness and row-volume anomalies
  • Alert with enough context to support action
Quality checks
assert no_nulls(RequestID)
assert unique(RequestID, UpdatedAt)
assert accepted_values(Status, ["Open","Closed","Canceled"])
assert freshness(max_age="2h")
✅

A green job is not enough; the data itself must also be healthy.

Practice — 3 cases

Practice 16 / Práctica 16
The pipeline ran successfully but today's records are missing. What signal is most relevant?
Practice 17 / Práctica 17
What makes an alert actionable?
Practice 18 / Práctica 18
Why keep lineage for the final KPI?
MODULE 07
🔐

Governance, Security & Platform

The platform assigns ownership, classifies sensitive fields, applies least privilege, protects secrets and defines retention. The deployment may be local, cloud or hybrid, but governance travels with the data.

✨
Key idea

Ownership → Classification → RBAC → Encryption → Audit → Retention

Core ideas

  • Classify PII and sensitive location/contact fields
  • Use service identities and least-privilege access
  • Keep secrets outside code and log access to governed assets
Policy sketch
AnalystRole: read curated analytics
PipelineRole: write curated tables
AdminRole: manage platform
PII: masked unless explicitly authorized
✅

Cloud does not replace governance; it changes where governance controls are implemented.

Practice — 3 cases

Practice 19 / Práctica 19
Who should normally see unmasked sensitive contact data?
Practice 20 / Práctica 20
Where should API credentials be stored?
Practice 21 / Práctica 21
What principle limits a pipeline identity to only required permissions?
MODULE 08
📊

Analytics, AI & Final Delivery

The curated pipeline now serves consumers. BI can expose service volume, backlog, resolution time and SLA performance. An API can deliver trusted metrics, while an AI assistant can retrieve governed definitions and explanatory context.

✨
Key idea

Trusted Data → Semantic Model → Dashboard / API / AI

Core ideas

  • Define KPIs once and reuse them across consumers
  • Expose only governed, purpose-fit data products
  • Use AI retrieval over trusted context rather than bypassing the data platform
Final architecture
Service Requests
   ↓
Ingest → Raw → Curated → Semantic Model
                     ↓
               BI / API / RAG
✅

The finished product is not the pipeline alone. It is a trusted decision system.

Practice — 3 cases

Practice 22 / Práctica 22
What should power the final KPI definitions?
Practice 23 / Práctica 23
What is the safest AI pattern for explaining service metrics?
Practice 24 / Práctica 24
What is the real deliverable of the capstone?
KNOWLEDGE CHECK

Rapid Review — Connect the Entire Pipeline

What is the first design input for this capstone?

Business outcome and consumer need

Why preserve a raw layer?

Replay, evidence and recovery

What should happen before downstream publishing?

Quality gates must pass

What principle protects sensitive data access?

Least privilege with explicit authorization

What is the final product?

A trusted end-to-end data product for decisions

CAPSTONE COMPLETE

From Separate Skills to One Trusted Data Product

The capstone turns the nine learning blocks into one operating system for trusted data.

Final architecture
Sources→Ingestion→Raw→Curated→Quality→Semantic→BI / API / AI

Certificate of Participation

Complete at least 12 of the 24 practices and enter your name to unlock the certificate.

Current participation: 0 / 24 • 0%

Capstone note

The architecture matters more than any single vendor tool.