Data Engineering for Humans • Training 04
Article-Training • Data Engineering for Humans

Processing & Transformation

Turn Raw Data into Trusted, Useful Data

Processing turns raw inputs into reliable data products that people and systems can actually use.

Raw Data → Clean → Standardize → Enrich → Model → Validate → Serve
Raw
→
Clean
→
Model
SQL / Python
+
dbt / Spark
8learning modules
24interactive practices
5rapid review questions
50%certificate threshold
Learning target
Build transformations that are correct, repeatable, scalable, and understandable.
Practice progress0 / 24

Complete 12 of 24 practices (50%) and enter your name to unlock the Certificate of Participation.

MODULE 01
⚙️

Transformation Foundations

Transformation changes the shape, quality, meaning, or usability of data so downstream consumers can trust it.

⚙️
Key idea

Cleaning removes obvious defects and inconsistencies • Standardization creates consistent formats and definitions • Enrichment and modeling add business context and analytical structure

Core ideas

  • Cleaning removes obvious defects and inconsistencies
  • Standardization creates consistent formats and definitions
  • Enrichment and modeling add business context and analytical structure
Conceptual model
Raw
  ↓
Clean + Standardize
  ↓
Enrich + Model
  ↓
Validated Data Product
✅

A transformation is successful when downstream users can rely on both the values and their meaning.

Practice — 3 cases

Practice 1 / Práctica 1
What is the core purpose of transformation?
Practice 2 / Práctica 2
Which step creates consistent date, code, or naming formats?
Practice 3 / Práctica 3
Which result is strongest?
MODULE 02
🧮

SQL Transformations

SQL is often the first transformation tool for structured data because joins, filters, aggregations, window functions, and set-based operations map naturally to relational workloads.

🧮
Key idea

Use set-based logic for relational transformations • Keep business rules explicit and reviewable • Break complex logic into understandable steps

Core ideas

  • Use set-based logic for relational transformations
  • Keep business rules explicit and reviewable
  • Break complex logic into understandable steps
Conceptual model
SELECT
  customer_id,
  SUM(amount) AS revenue
FROM sales
WHERE status = 'Posted'
GROUP BY customer_id;
✅

Prefer SQL when the data is relational and the transformation can be expressed clearly with set-based operations.

Practice — 3 cases

Practice 4 / Práctica 4
Which workload is a strong fit for SQL transformation?
Practice 5 / Práctica 5
Why are set-based SQL operations valuable?
Practice 6 / Práctica 6
What improves maintainability most?
MODULE 03
🐍

Python & Pandas

Python adds flexibility for file handling, APIs, irregular business logic, data profiling, automation, and transformations that are awkward to express in SQL alone.

🐍
Key idea

Pandas is convenient for tabular transformations in memory • Python integrates files, APIs, validation, and automation • Large datasets may require engines beyond a single-machine DataFrame

Core ideas

  • Pandas is convenient for tabular transformations in memory
  • Python integrates files, APIs, validation, and automation
  • Large datasets may require engines beyond a single-machine DataFrame
Conceptual model
df = (df
  .drop_duplicates()
  .assign(total=lambda x: x.qty * x.price)
  .query("status == 'Active'"))
✅

Use Python where flexibility matters, but keep transformations deterministic and testable.

Practice — 3 cases

Practice 7 / Práctica 7
When is Python especially useful in a data pipeline?
Practice 8 / Práctica 8
What does Pandas primarily provide?
Practice 9 / Práctica 9
What is a sensible limit to remember?
MODULE 04
🧱

dbt & Modular SQL

dbt organizes SQL transformations into modular models with dependencies, tests, documentation, and repeatable builds inside analytical data platforms.

🧱
Key idea

Models make SQL transformation logic modular • Dependencies form a directed transformation graph • Tests and documentation move quality closer to transformation code

Core ideas

  • Models make SQL transformation logic modular
  • Dependencies form a directed transformation graph
  • Tests and documentation move quality closer to transformation code
Conceptual model
source → staging → intermediate → mart
          ↘ tests ↗
       documented lineage
✅

dbt is strongest when transformation logic belongs in SQL and you want software-engineering discipline around it.

Practice — 3 cases

Practice 10 / Práctica 10
What problem does dbt primarily address?
Practice 11 / Práctica 11
What does a dbt dependency graph represent?
Practice 12 / Práctica 12
Which practice best fits dbt?
MODULE 05
✨

Spark & Distributed Processing

Spark distributes data processing across a cluster so transformations can scale beyond the memory and compute of one machine.

✨
Key idea

Distributed partitions allow parallel processing • Shuffles can be expensive and should be understood • Spark is useful when scale or workload complexity exceeds one machine

Core ideas

  • Distributed partitions allow parallel processing
  • Shuffles can be expensive and should be understood
  • Spark is useful when scale or workload complexity exceeds one machine
Conceptual model
Dataset
  ↓ partition
Workers → Transform in parallel
  ↓ shuffle if needed
Unified Result
✅

Do not choose Spark just because it is powerful; choose it when distributed processing solves a real scale problem.

Practice — 3 cases

Practice 13 / Práctica 13
Why use Spark?
Practice 14 / Práctica 14
What is a shuffle?
Practice 15 / Práctica 15
When is Spark least justified?
MODULE 06
⏱️

Batch vs Streaming

Batch processes bounded groups of data on a schedule; streaming processes events continuously or near real time. The business latency requirement should drive the choice.

⏱️
Key idea

Batch is simpler for scheduled, bounded workloads • Streaming reduces latency but adds state and operational complexity • Many platforms combine batch and streaming patterns

Core ideas

  • Batch is simpler for scheduled, bounded workloads
  • Streaming reduces latency but adds state and operational complexity
  • Many platforms combine batch and streaming patterns
Conceptual model
Batch:  [events] → schedule → transform
Stream: event → event → event → continuous transform
✅

Use the simplest processing mode that meets the business freshness requirement.

Practice — 3 cases

Practice 16 / Práctica 16
What defines batch processing?
Practice 17 / Práctica 17
Why choose streaming?
Practice 18 / Práctica 18
What is the best design rule?
MODULE 07
🧪

Safe Transformation Patterns

Production transformations should be deterministic, idempotent where possible, observable, versioned, and protected by validation so reruns do not corrupt results.

🧪
Key idea

Idempotency makes reruns safer • Validation catches schema and business-rule violations • Lineage and versioning make changes explainable

Core ideas

  • Idempotency makes reruns safer
  • Validation catches schema and business-rule violations
  • Lineage and versioning make changes explainable
Conceptual model
Input
 ↓
Transform → Validate → Publish
    ↘ logs + metrics + lineage
Rerun safely when needed
✅

Transformation logic is production software: test it, observe it, and design for controlled reruns.

Practice — 3 cases

Practice 19 / Práctica 19
What does idempotent transformation mean?
Practice 20 / Práctica 20
Why validate after transformation?
Practice 21 / Práctica 21
Which combination is strongest?
MODULE 08
🧭

Choosing the Right Processing Engine

The best engine is the simplest one that satisfies data volume, latency, transformation complexity, ecosystem, team skills, operational cost, and reliability needs.

🧭
Key idea

SQL fits relational set-based transformations • Python fits flexible automation and custom logic • Spark fits distributed scale; dbt fits modular analytical SQL

Core ideas

  • SQL fits relational set-based transformations
  • Python fits flexible automation and custom logic
  • Spark fits distributed scale; dbt fits modular analytical SQL
Conceptual model
Need
 ├─ Relational → SQL / dbt
 ├─ Flexible logic → Python
 ├─ Distributed scale → Spark
 └─ Low latency → Streaming engine
✅

Architecture improves when tool selection follows workload requirements instead of tool popularity.

Practice — 3 cases

Practice 22 / Práctica 22
What should drive processing-engine selection?
Practice 23 / Práctica 23
Which pairing is most natural?
Practice 24 / Práctica 24
What is the mature architectural choice?

5-Question Knowledge Check

Can you explain how raw data becomes a reliable data product?

Transformation changes data for a purpose.

Cleaning, standardization, enrichment, modeling, and validation make data usable while preserving business meaning.

SQL and Python solve different transformation problems.

SQL excels at relational set operations; Python adds flexible automation, file/API handling, and custom logic.

dbt brings engineering discipline to analytical SQL.

Modular models, dependency graphs, tests, documentation, and lineage make transformations easier to maintain.

Spark solves distributed-scale processing problems.

Spark partitions work across multiple workers, but distributed shuffles and operations add complexity and cost.

Latency should drive batch vs streaming.

Use batch when scheduled processing is sufficient; use streaming when the business genuinely needs low-latency event handling.

Processing Decision Map

Choose the simplest transformation engine that meets the workload requirement.

NeedStrong candidate
Relational joins, filters, aggregatesSQL
Files, APIs, custom logic, automationPython / Pandas
Modular analytical SQL with tests and lineagedbt
Distributed processing at large scaleSpark
Scheduled bounded processingBatch
Low-latency event processingStreaming

Certificate of Participation

Complete at least 12 of the 24 practice cases (50%) and enter your name.

0 / 24 • 0%

Production note

Reliable transformation is about repeatable meaning, not one successful run.