Data Engineering for Humans • Training 03
Article-Training • Data Engineering for Humans

Storage & Databases

Choose the Right Place for Every Kind of Data

Good storage design matches the workload instead of forcing every workload into the same system.

Workload → Structure → Scale → Storage Pattern → Performance & Cost
Sources
→
Ingest
→
Land
Batch
|
Streaming
|
CDC
8learning modules
24interactive practices
5rapid review questions
50%certificate threshold
Learning target
Choose storage architecture based on workload, data shape and access pattern.
Practice progress0 / 24

Complete 12 of 24 practices (50%) and enter your name to unlock the Certificate of Participation.

MODULE 01
🗄️

Storage Foundations

Storage architecture decides where data should live based on workload, structure, latency, scale, durability, access patterns, governance, and cost.

🗄️
Key idea

Different workloads need different storage patterns • Operational and analytical systems optimize for different behaviors • Durability, availability, latency, scale, and cost must be balanced

Core ideas

  • Different workloads need different storage patterns
  • Operational and analytical systems optimize for different behaviors
  • Durability, availability, latency, scale, and cost must be balanced
Conceptual model
Workload
   ↓
Access Pattern
   ↓
Structure + Scale
   ↓
Storage Choice
   ↓
Performance + Cost
✅

Start with the workload and access pattern—not with a favorite database product.

Practice — 3 cases

Practice 1 / Práctica 1
What is the main goal of a storage architecture?
Practice 2 / Práctica 2
Which characteristic most clearly distinguishes storage choices?
Practice 3 / Práctica 3
Which principle is strongest?
MODULE 02
⚖️

OLTP vs OLAP

OLTP systems optimize short, frequent transactions. OLAP systems optimize large analytical scans, aggregations, trends, and historical analysis.

⚖️
Key idea

OLTP favors fast inserts, updates, and point lookups • OLAP favors large reads, joins, aggregations, and history • Separate operational and analytical workloads when their needs conflict

Core ideas

  • OLTP favors fast inserts, updates, and point lookups
  • OLAP favors large reads, joins, aggregations, and history
  • Separate operational and analytical workloads when their needs conflict
Conceptual model
OLTP → Orders / Payments / Updates
OLAP → Trends / Aggregations / History
✅

Do not force heavy analytics onto a transactional database when it threatens operational performance.

Practice — 3 cases

Practice 4 / Práctica 4
Which workload is most associated with OLTP?
Practice 5 / Práctica 5
Which workload is most associated with OLAP?
Practice 6 / Práctica 6
Why separate OLTP and OLAP workloads?
MODULE 03
🏢

Data Warehouse

A data warehouse stores curated, integrated, historical data optimized for analytics, reporting, KPI calculation, and repeatable business definitions.

🏢
Key idea

Curated and modeled data supports consistent analytics • Historical data enables trends and comparisons • Warehouses prioritize analytical performance over row-by-row transactions

Core ideas

  • Curated and modeled data supports consistent analytics
  • Historical data enables trends and comparisons
  • Warehouses prioritize analytical performance over row-by-row transactions
Conceptual model
Operational Sources
   ↓
Transform / Model
   ↓
Data Warehouse
   ↓
BI / Reporting
✅

Use a warehouse when trusted, repeatable analytical structure matters more than raw-data flexibility.

Practice — 3 cases

Practice 7 / Práctica 7
What is a data warehouse primarily optimized for?
Practice 8 / Práctica 8
Why is historical data valuable in a warehouse?
Practice 9 / Práctica 9
Which consumer commonly uses warehouse data?
MODULE 04
🌊

Data Lake

A data lake stores large volumes of raw or lightly processed structured, semi-structured, and unstructured data, usually on scalable object storage.

🌊
Key idea

Lakes preserve raw data for future uses • Schema-on-read allows interpretation at consumption time • Without governance, a lake can become a data swamp

Core ideas

  • Lakes preserve raw data for future uses
  • Schema-on-read allows interpretation at consumption time
  • Without governance, a lake can become a data swamp
Conceptual model
Files + Events + Tables
   ↓
Object Storage / Data Lake
   ↓
SQL / Spark / ML / Exploration
✅

A lake is flexible, but flexibility still requires metadata, lineage, access control, and lifecycle management.

Practice — 3 cases

Practice 10 / Práctica 10
How is a data lake best described?
Practice 11 / Práctica 11
What does schema-on-read mean?
Practice 12 / Práctica 12
What is a risk of a poorly governed data lake?
MODULE 05
🏠

Lakehouse

A lakehouse combines low-cost flexible lake storage with warehouse-like table management, reliability, governance, and analytical performance.

🏠
Key idea

Open or columnar files can be managed as reliable tables • Transaction layers add consistency and table semantics • One platform can support BI, data science, and ML over shared data

Core ideas

  • Open or columnar files can be managed as reliable tables
  • Transaction layers add consistency and table semantics
  • One platform can support BI, data science, and ML over shared data
Conceptual model
Object Storage
   +
Table / Transaction Layer
   ↓
Lakehouse
   ↓
BI + SQL + ML
✅

A lakehouse is useful when you want lake flexibility without giving up reliable analytical table behavior.

Practice — 3 cases

Practice 13 / Práctica 13
What is the core lakehouse idea?
Practice 14 / Práctica 14
What capability does a lakehouse add over a raw lake?
Practice 15 / Práctica 15
Who can benefit from shared lakehouse data?
MODULE 06
🧩

SQL vs NoSQL

Relational SQL databases excel when relationships, schema, transactions, and consistency matter. NoSQL databases trade some relational structure for flexible models, scale patterns, or specialized access.

🧩
Key idea

SQL is strong for structured relationships and ACID transactions • NoSQL includes document, key-value, wide-column, and graph patterns • Choose based on data model and access pattern, not fashion

Core ideas

  • SQL is strong for structured relationships and ACID transactions
  • NoSQL includes document, key-value, wide-column, and graph patterns
  • Choose based on data model and access pattern, not fashion
Conceptual model
Relational Need → SQL
Flexible / Specialized Access → NoSQL
Hybrid Workload → Polyglot Persistence
✅

SQL and NoSQL are not enemies; mature architectures often use each where it fits best.

Practice — 3 cases

Practice 16 / Práctica 16
Why choose a relational SQL database?
Practice 17 / Práctica 17
When can NoSQL be a strong fit?
Practice 18 / Práctica 18
What should drive SQL vs NoSQL selection?
MODULE 07
📦

File & Object Storage

File and object storage are foundational for datasets, backups, logs, media, model artifacts, and lake architectures. Formats and partitioning strongly affect analytical performance.

📦
Key idea

Object storage scales well for large collections of files and blobs • Columnar formats such as Parquet reduce scan and storage cost for analytics • Partitioning helps engines skip irrelevant data

Core ideas

  • Object storage scales well for large collections of files and blobs
  • Columnar formats such as Parquet reduce scan and storage cost for analytics
  • Partitioning helps engines skip irrelevant data
Conceptual model
Object Storage
  /year=2026/month=09/
     part-001.parquet
     part-002.parquet
✅

For analytical files, format and partition design can matter almost as much as where the files are stored.

Practice — 3 cases

Practice 19 / Práctica 19
What is object storage especially good for?
Practice 20 / Práctica 20
Why can Parquet improve analytics?
Practice 21 / Práctica 21
Why does partitioning matter in large datasets?
MODULE 08
🎯

Choosing the Right Storage

The final storage decision balances performance, latency, durability, availability, retention, governance, cost, interoperability, and operational complexity.

🎯
Key idea

Hot data needs fast access; cold data can use cheaper tiers • Retention follows business, legal, and operational requirements • Hybrid architectures are normal when workloads differ

Core ideas

  • Hot data needs fast access; cold data can use cheaper tiers
  • Retention follows business, legal, and operational requirements
  • Hybrid architectures are normal when workloads differ
Conceptual model
Transactions → Operational DB
Analytics → Warehouse / Lakehouse
Raw / Archive → Object Storage
Specialized Access → NoSQL
✅

The right architecture is the least-complex combination that meets reliability, performance, governance, and cost requirements.

Practice — 3 cases

Practice 22 / Práctica 22
What should guide storage selection first?
Practice 23 / Práctica 23
What is the purpose of storage tiers?
Practice 24 / Práctica 24
What does durability describe?

5-Question Knowledge Check

Can you explain why storage design changes with workload?

OLTP and OLAP optimize different workloads.

OLTP favors short transactions and point operations; OLAP favors scans, joins, aggregations, and historical analysis.

A warehouse is curated for analytics.

Warehouses integrate and model trusted historical data for reporting, KPIs, and analytical queries.

A lake preserves flexible raw data.

Lakes support structured, semi-structured, and unstructured data.

A lakehouse adds reliable table behavior to lake storage.

Lakehouse designs combine flexibility with table reliability.

Storage selection starts with workload and access pattern.

Latency, consistency, scale, durability, retention, governance, cost, and operational complexity refine the final choice.

Storage Decision Map

Choose storage by workload, access pattern, structure, latency, and governance.

NeedStrong candidate
High-volume operational transactionsRelational OLTP database
Curated historical analyticsData Warehouse
Raw / semi-structured / unstructured data at scaleData Lake / Object Storage
Lake flexibility + managed analytical tablesLakehouse
Flexible or specialized non-relational accessNoSQL
Low-cost archive / large files / model artifactsObject Storage

Certificate of Participation

Complete at least 12 of the 24 practice cases (50%) and enter your name.

0 / 24 • 0%

Production note

Good storage architecture matches the workload, data shape, scale, latency, durability, governance, and access pattern instead of forcing every use case into one system.