Ingestion is how data enters a platform. Integration is how different systems, formats, and events are connected so data can move reliably from source to destination. This training focuses on the patterns that make that movement repeatable, observable, and safe.
Complete 12 of 24 practices (50%) and enter your name to unlock the Certificate of Participation.
Ingestion brings data into a platform. Integration connects systems so data can move, synchronize, and remain useful across boundaries.
Ingestion answers “How does data enter?” • Integration answers “How do systems exchange and coordinate data?” • Good pipelines define source, transport, destination, and failure behavior
Source System
↓
Ingestion
↓
Landing / Staging
↓
Integration / Transformation
↓
DestinationDefine the movement contract before choosing a tool: source, frequency, latency, schema, destination, and recovery path.
APIs, managed connectors, and files are common ways to ingest operational data. Each method trades control, convenience, latency, and maintenance differently.
APIs offer programmable access • Connectors reduce custom code • Files remain common for scheduled exchange • Authentication and pagination are first-class concerns
API → authenticate → request → paginate → land
Connector → configure → sync → monitor
File → detect → validate → load → archiveChoose the simplest interface that still gives you the reliability, freshness, and control the business needs.
Batch ingestion moves groups of records on a schedule or when a file or dataset becomes available. It is simple, efficient, and often the right choice when low latency is not required.
Batch works well for hourly, nightly, or periodic loads • Watermarks track what has already been loaded • Late-arriving data and restartability must be designed
Last Watermark = 10:00
↓
Read 10:00–11:00
↓
Validate + Load
↓
Commit Watermark = 11:00If the business can tolerate minutes or hours of latency, batch is often cheaper and simpler than streaming.
Streaming ingestion processes events continuously or near continuously. Message queues and event brokers decouple producers from consumers and absorb bursts of activity.
Producers publish events • Brokers buffer and route them • Consumers process them independently • Ordering, offsets, and delivery guarantees matter
Producer → Broker / Queue → Consumer
↓
Replay / Offset
↓
Processing StateUse streaming when latency has business value—not merely because the technology is available.
A webhook lets one system notify another when something happens. Instead of polling repeatedly, the receiver exposes an endpoint and reacts to incoming events.
Webhooks are push-based • Receivers must authenticate, validate, and respond quickly • Durable processing should happen after safe receipt
Event occurs
↓
Sender POSTs webhook
↓
Verify + acknowledge
↓
Queue / persist
↓
Process safelyDesign webhook receivers for duplicate delivery and temporary outages; external senders may retry.
CDC captures inserts, updates, and deletes from a source without repeatedly re-reading the entire dataset. It is a key pattern for keeping downstream systems synchronized efficiently.
CDC focuses on changes rather than full reloads • Ordering and transaction boundaries matter • Downstream consumers must handle inserts, updates, and deletes correctly
Source DB
├─ INSERT
├─ UPDATE
└─ DELETE
↓
CDC Log
↓
Checkpoint → Consumer → TargetCDC is powerful for incremental synchronization, but correctness depends on preserving change order and recovery position.
Reliable ingestion assumes things will fail: networks time out, APIs throttle, events repeat, and schemas evolve. Retry policies, idempotency, dead-letter handling, and schema validation make pipelines resilient.
Retry transient failures with limits and backoff • Idempotent loads make repeats safe • Schema drift must be detected before it silently corrupts downstream logic
Request fails
↓
Retry + backoff
↓
Still failing? → Dead-letter / alert
Duplicate input → Idempotent load → Same final stateReliability is not “never failing”; it is failing predictably, recovering safely, and proving what happened.
Incoming data should land in a controlled area before becoming trusted. A landing or staging zone preserves raw evidence, supports validation, and separates ingestion from downstream transformation.
Land raw data before destructive transformation • Validate counts, schema, freshness, and basic rules • Protect secrets and sensitive fields • Hand off only after acceptance checks pass
Source
↓
Secure Transport
↓
Landing / Raw
↓
Validate
↓
Accept → Transform / Publish
Reject → Quarantine / AlertTreat ingestion as a controlled handoff: receive, preserve, validate, secure, and only then promote data downstream.
Can you design reliable movement from source to landing zone? Open each item after answering it in your own words. The 24 interactive practices above drive certificate progress.
Ingestion brings data into a platform; integration connects systems so data can move, synchronize, and remain useful across boundaries.
When minutes or hours of latency are acceptable and scheduled grouped processing is simpler and more economical.
It captures source inserts, updates, and deletes incrementally so downstream systems can stay synchronized without full reloads.
Because retries and duplicate deliveries happen; idempotency lets the same input be processed again without unintended duplicate effects.
To preserve incoming data, validate it, support recovery and traceability, and separate receipt from downstream transformation.
Choose the pattern based on how data arrives, how fast it must move, and how safely it must recover.
| Need | Strong candidate |
|---|---|
| Scheduled file or periodic extract | Batch ingestion |
| Programmatic source access | API / Connector |
| Push notification when an event occurs | Webhook |
| Continuous low-latency events | Streaming / Message Queue |
| Incremental database changes | CDC |
| Safe receipt before transformation | Landing Zone + Validation |
Complete at least 12 of the 24 practice cases (50%) and enter your name.
Reliable ingestion is not about moving data once; it is about moving it repeatedly with traceability, validation, recovery, and controlled failure behavior.