Data Engineering for Humans • Training 02
Article-Training • Data Engineering for Humans

Ingestion & Integration

Move Data In Reliably — APIs, Batch, Streaming, CDC & Connectors

Ingestion is how data enters a platform. Integration is how different systems, formats, and events are connected so data can move reliably from source to destination. This training focuses on the patterns that make that movement repeatable, observable, and safe.

Sources → Connectors / APIs → Batch or Streaming → CDC / Queues → Landing Zone → Validated Data
Sources
→
Ingest
→
Land
Batch
|
Streaming
|
CDC
8learning modules
24interactive practices
5rapid review questions
50%certificate threshold
Learning target
Design reliable ingestion and integration flows using APIs, connectors, batch, streaming, message queues, webhooks, and change data capture while handling retries, duplicates, schema changes, and security.
Practice progress0 / 24

Complete 12 of 24 practices (50%) and enter your name to unlock the Certificate of Participation.

MODULE 01
📥

Ingestion vs Integration

Ingestion brings data into a platform. Integration connects systems so data can move, synchronize, and remain useful across boundaries.

📥
Key idea

Ingestion answers “How does data enter?” • Integration answers “How do systems exchange and coordinate data?” • Good pipelines define source, transport, destination, and failure behavior

Core ideas

  • Ingestion moves data into a platform
  • Integration connects systems and data flows
  • Reliable movement requires explicit failure and retry behavior
Conceptual model
Source System
   ↓
Ingestion
   ↓
Landing / Staging
   ↓
Integration / Transformation
   ↓
Destination
✅

Define the movement contract before choosing a tool: source, frequency, latency, schema, destination, and recovery path.

Practice — 3 cases

Practice 1 / Práctica 1
What does ingestion primarily do?
Practice 2 / Práctica 2
What best describes integration?
Practice 3 / Práctica 3
Which item belongs in a movement contract?
MODULE 02
🔌

APIs, Connectors & File-Based Ingestion

APIs, managed connectors, and files are common ways to ingest operational data. Each method trades control, convenience, latency, and maintenance differently.

🔌
Key idea

APIs offer programmable access • Connectors reduce custom code • Files remain common for scheduled exchange • Authentication and pagination are first-class concerns

Core ideas

  • APIs expose data through programmable endpoints
  • Connectors package common integration logic
  • File ingestion needs naming, location, format, and arrival rules
Conceptual model
API → authenticate → request → paginate → land
Connector → configure → sync → monitor
File → detect → validate → load → archive
✅

Choose the simplest interface that still gives you the reliability, freshness, and control the business needs.

Practice — 3 cases

Practice 4 / Práctica 4
Why are managed connectors useful?
Practice 5 / Práctica 5
What must an API ingestion process often handle?
Practice 6 / Práctica 6
What is important for file ingestion?
MODULE 03
🕒

Batch Ingestion

Batch ingestion moves groups of records on a schedule or when a file or dataset becomes available. It is simple, efficient, and often the right choice when low latency is not required.

🕒
Key idea

Batch works well for hourly, nightly, or periodic loads • Watermarks track what has already been loaded • Late-arriving data and restartability must be designed

Core ideas

  • Schedules or file arrival can trigger a batch
  • Watermarks help load only new or changed ranges
  • Restartable batches avoid full reloads after failure
Conceptual model
Last Watermark = 10:00
   ↓
Read 10:00–11:00
   ↓
Validate + Load
   ↓
Commit Watermark = 11:00
✅

If the business can tolerate minutes or hours of latency, batch is often cheaper and simpler than streaming.

Practice — 3 cases

Practice 7 / Práctica 7
When is batch ingestion a strong fit?
Practice 8 / Práctica 8
What does a watermark usually track?
Practice 9 / Práctica 9
Why design batches to be restartable?
MODULE 04
⚡

Streaming, Events & Message Queues

Streaming ingestion processes events continuously or near continuously. Message queues and event brokers decouple producers from consumers and absorb bursts of activity.

⚡
Key idea

Producers publish events • Brokers buffer and route them • Consumers process them independently • Ordering, offsets, and delivery guarantees matter

Core ideas

  • Streaming is driven by events and low latency
  • Queues decouple producers from consumers
  • Consumers need a strategy for duplicates and replay
Conceptual model
Producer → Broker / Queue → Consumer
             ↓
          Replay / Offset
             ↓
        Processing State
✅

Use streaming when latency has business value—not merely because the technology is available.

Practice — 3 cases

Practice 10 / Práctica 10
What is a major benefit of a message queue?
Practice 11 / Práctica 11
What does a consumer do?
Practice 12 / Práctica 12
When is streaming most justified?
MODULE 05
🪝

Webhooks & Event-Driven Integration

A webhook lets one system notify another when something happens. Instead of polling repeatedly, the receiver exposes an endpoint and reacts to incoming events.

🪝
Key idea

Webhooks are push-based • Receivers must authenticate, validate, and respond quickly • Durable processing should happen after safe receipt

Core ideas

  • Push avoids unnecessary polling
  • Signatures or secrets help verify the sender
  • Acknowledge first, then process reliably when possible
Conceptual model
Event occurs
   ↓
Sender POSTs webhook
   ↓
Verify + acknowledge
   ↓
Queue / persist
   ↓
Process safely
✅

Design webhook receivers for duplicate delivery and temporary outages; external senders may retry.

Practice — 3 cases

Practice 13 / Práctica 13
How is a webhook different from polling?
Practice 14 / Práctica 14
Why verify a webhook signature or secret?
Practice 15 / Práctica 15
What should a robust webhook design expect?
MODULE 06
🔄

Change Data Capture (CDC)

CDC captures inserts, updates, and deletes from a source without repeatedly re-reading the entire dataset. It is a key pattern for keeping downstream systems synchronized efficiently.

🔄
Key idea

CDC focuses on changes rather than full reloads • Ordering and transaction boundaries matter • Downstream consumers must handle inserts, updates, and deletes correctly

Core ideas

  • CDC reduces unnecessary full-table scans
  • Deletes are part of the change stream too
  • Checkpoints let consumers resume from a known position
Conceptual model
Source DB
  ├─ INSERT
  ├─ UPDATE
  └─ DELETE
      ↓
   CDC Log
      ↓
Checkpoint → Consumer → Target
✅

CDC is powerful for incremental synchronization, but correctness depends on preserving change order and recovery position.

Practice — 3 cases

Practice 16 / Práctica 16
What does CDC capture?
Practice 17 / Práctica 17
Why use a CDC checkpoint?
Practice 18 / Práctica 18
What is an important CDC design concern?
MODULE 07
🧯

Retries, Idempotency & Schema Drift

Reliable ingestion assumes things will fail: networks time out, APIs throttle, events repeat, and schemas evolve. Retry policies, idempotency, dead-letter handling, and schema validation make pipelines resilient.

🧯
Key idea

Retry transient failures with limits and backoff • Idempotent loads make repeats safe • Schema drift must be detected before it silently corrupts downstream logic

Core ideas

  • Retries should be bounded and observable
  • Idempotency means the same input can be processed again safely
  • Schema validation catches incompatible structural changes
Conceptual model
Request fails
   ↓
Retry + backoff
   ↓
Still failing? → Dead-letter / alert

Duplicate input → Idempotent load → Same final state
✅

Reliability is not “never failing”; it is failing predictably, recovering safely, and proving what happened.

Practice — 3 cases

Practice 19 / Práctica 19
What does idempotent ingestion provide?
Practice 20 / Práctica 20
What is schema drift?
Practice 21 / Práctica 21
What is a good retry practice?
MODULE 08
🛬

Landing Zones, Validation & Secure Handoff

Incoming data should land in a controlled area before becoming trusted. A landing or staging zone preserves raw evidence, supports validation, and separates ingestion from downstream transformation.

🛬
Key idea

Land raw data before destructive transformation • Validate counts, schema, freshness, and basic rules • Protect secrets and sensitive fields • Hand off only after acceptance checks pass

Core ideas

  • Landing zones preserve the original delivery
  • Validation establishes whether the load is acceptable
  • Secure handoff protects credentials and sensitive data
Conceptual model
Source
  ↓
Secure Transport
  ↓
Landing / Raw
  ↓
Validate
  ↓
Accept → Transform / Publish
Reject → Quarantine / Alert
✅

Treat ingestion as a controlled handoff: receive, preserve, validate, secure, and only then promote data downstream.

Practice — 3 cases

Practice 22 / Práctica 22
Why use a landing zone?
Practice 23 / Práctica 23
What should happen before promoting a load downstream?
Practice 24 / Práctica 24
What is the best final state for ingestion?

5-Question Knowledge Check

Can you design reliable movement from source to landing zone? Open each item after answering it in your own words. The 24 interactive practices above drive certificate progress.

1. What is the difference between ingestion and integration?

Ingestion brings data into a platform; integration connects systems so data can move, synchronize, and remain useful across boundaries.

2. When is batch preferable to streaming?

When minutes or hours of latency are acceptable and scheduled grouped processing is simpler and more economical.

3. What problem does CDC solve?

It captures source inserts, updates, and deletes incrementally so downstream systems can stay synchronized without full reloads.

4. Why does idempotency matter?

Because retries and duplicate deliveries happen; idempotency lets the same input be processed again without unintended duplicate effects.

5. Why use a landing zone?

To preserve incoming data, validate it, support recovery and traceability, and separate receipt from downstream transformation.

Ingestion & Integration Map

Choose the pattern based on how data arrives, how fast it must move, and how safely it must recover.

NeedStrong candidate
Scheduled file or periodic extractBatch ingestion
Programmatic source accessAPI / Connector
Push notification when an event occursWebhook
Continuous low-latency eventsStreaming / Message Queue
Incremental database changesCDC
Safe receipt before transformationLanding Zone + Validation

Certificate of Participation

Complete at least 12 of the 24 practice cases (50%) and enter your name.

0 / 24 • 0%

Production note

Reliable ingestion is not about moving data once; it is about moving it repeatedly with traceability, validation, recovery, and controlled failure behavior.