Good storage design matches the workload instead of forcing every workload into the same system.
Complete 12 of 24 practices (50%) and enter your name to unlock the Certificate of Participation.
Storage architecture decides where data should live based on workload, structure, latency, scale, durability, access patterns, governance, and cost.
Different workloads need different storage patterns • Operational and analytical systems optimize for different behaviors • Durability, availability, latency, scale, and cost must be balanced
Workload
↓
Access Pattern
↓
Structure + Scale
↓
Storage Choice
↓
Performance + CostStart with the workload and access pattern—not with a favorite database product.
OLTP systems optimize short, frequent transactions. OLAP systems optimize large analytical scans, aggregations, trends, and historical analysis.
OLTP favors fast inserts, updates, and point lookups • OLAP favors large reads, joins, aggregations, and history • Separate operational and analytical workloads when their needs conflict
OLTP → Orders / Payments / Updates
OLAP → Trends / Aggregations / HistoryDo not force heavy analytics onto a transactional database when it threatens operational performance.
A data warehouse stores curated, integrated, historical data optimized for analytics, reporting, KPI calculation, and repeatable business definitions.
Curated and modeled data supports consistent analytics • Historical data enables trends and comparisons • Warehouses prioritize analytical performance over row-by-row transactions
Operational Sources
↓
Transform / Model
↓
Data Warehouse
↓
BI / ReportingUse a warehouse when trusted, repeatable analytical structure matters more than raw-data flexibility.
A data lake stores large volumes of raw or lightly processed structured, semi-structured, and unstructured data, usually on scalable object storage.
Lakes preserve raw data for future uses • Schema-on-read allows interpretation at consumption time • Without governance, a lake can become a data swamp
Files + Events + Tables
↓
Object Storage / Data Lake
↓
SQL / Spark / ML / ExplorationA lake is flexible, but flexibility still requires metadata, lineage, access control, and lifecycle management.
A lakehouse combines low-cost flexible lake storage with warehouse-like table management, reliability, governance, and analytical performance.
Open or columnar files can be managed as reliable tables • Transaction layers add consistency and table semantics • One platform can support BI, data science, and ML over shared data
Object Storage
+
Table / Transaction Layer
↓
Lakehouse
↓
BI + SQL + MLA lakehouse is useful when you want lake flexibility without giving up reliable analytical table behavior.
Relational SQL databases excel when relationships, schema, transactions, and consistency matter. NoSQL databases trade some relational structure for flexible models, scale patterns, or specialized access.
SQL is strong for structured relationships and ACID transactions • NoSQL includes document, key-value, wide-column, and graph patterns • Choose based on data model and access pattern, not fashion
Relational Need → SQL
Flexible / Specialized Access → NoSQL
Hybrid Workload → Polyglot PersistenceSQL and NoSQL are not enemies; mature architectures often use each where it fits best.
File and object storage are foundational for datasets, backups, logs, media, model artifacts, and lake architectures. Formats and partitioning strongly affect analytical performance.
Object storage scales well for large collections of files and blobs • Columnar formats such as Parquet reduce scan and storage cost for analytics • Partitioning helps engines skip irrelevant data
Object Storage
/year=2026/month=09/
part-001.parquet
part-002.parquetFor analytical files, format and partition design can matter almost as much as where the files are stored.
The final storage decision balances performance, latency, durability, availability, retention, governance, cost, interoperability, and operational complexity.
Hot data needs fast access; cold data can use cheaper tiers • Retention follows business, legal, and operational requirements • Hybrid architectures are normal when workloads differ
Transactions → Operational DB
Analytics → Warehouse / Lakehouse
Raw / Archive → Object Storage
Specialized Access → NoSQLThe right architecture is the least-complex combination that meets reliability, performance, governance, and cost requirements.
Can you explain why storage design changes with workload?
OLTP favors short transactions and point operations; OLAP favors scans, joins, aggregations, and historical analysis.
Warehouses integrate and model trusted historical data for reporting, KPIs, and analytical queries.
Lakes support structured, semi-structured, and unstructured data.
Lakehouse designs combine flexibility with table reliability.
Latency, consistency, scale, durability, retention, governance, cost, and operational complexity refine the final choice.
Choose storage by workload, access pattern, structure, latency, and governance.
| Need | Strong candidate |
|---|---|
| High-volume operational transactions | Relational OLTP database |
| Curated historical analytics | Data Warehouse |
| Raw / semi-structured / unstructured data at scale | Data Lake / Object Storage |
| Lake flexibility + managed analytical tables | Lakehouse |
| Flexible or specialized non-relational access | NoSQL |
| Low-cost archive / large files / model artifacts | Object Storage |
Complete at least 12 of the 24 practice cases (50%) and enter your name.
Good storage architecture matches the workload, data shape, scale, latency, durability, governance, and access pattern instead of forcing every use case into one system.