Data Scientist → AI Engineer • Training 04
Article-Training • AI Engineering Foundations

Cloud & Deployment

From Local Container to Reliable Production Service

Learn the deployment mindset: choose the right compute, externalize configuration, release safely, observe production behavior and scale AI workloads without losing control.

Image → Registry → Compute → Network → Observe → Scale → Improve
📦 IMAGE
→
☁️ CLOUD
→
📡 SERVICE
→
📈 SCALE
8learning modules
24interactive practices
5rapid review questions
50%certificate threshold
Learning target
Understand core cloud deployment choices, environments, managed services, networking, CI/CD, observability, scaling, reliability and cost tradeoffs for AI APIs and applications.
Practice progress0 / 24

Complete 12 of 24 practices (50%) and enter your name to unlock the Certificate of Participation.

MODULE 01
🧭

Cloud Mental Model

Cloud deployment means renting managed compute, storage, networking and platform capabilities instead of owning every physical layer. The engineering job is still to define boundaries and reliability expectations.

👁️
See it this way

Start with the simplest managed option that meets requirements. Complexity has an operational cost.

Core ideas

  • Compute runs code
  • Storage keeps durable data
  • Networking connects services
  • Managed services reduce operational burden
Try this
Local: container on laptop
Cloud: container on managed compute

Same app, different operational environment
✅

Start with the simplest managed option that meets requirements. Complexity has an operational cost.

Practice — 3 cases

Practice 1 / Práctica 1
Which statement best reflects this module?
Practice 2 / Práctica 2
Which action is the best engineering choice?
Practice 3 / Práctica 3
Which option would you avoid in a production system?
MODULE 02
🧱

Choose a Deployment Target

AI services can run on virtual machines, managed container platforms, serverless functions or orchestrated clusters. The right target depends on traffic, runtime, latency, GPU needs and team maturity.

👁️
See it this way

Do not choose Kubernetes because it sounds advanced. Choose it when the operational problem justifies it.

Core ideas

  • VM: maximum control
  • Managed containers: balanced default
  • Serverless: event/short workloads
  • Cluster/Kubernetes: complex scale/control
Try this
# Decision questions
# - long-running API?
# - bursty traffic?
# - GPU?
# - custom networking?
# - team ops capacity?
✅

Do not choose Kubernetes because it sounds advanced. Choose it when the operational problem justifies it.

Practice — 3 cases

Practice 4 / Práctica 4
Which statement best reflects this module?
Practice 5 / Práctica 5
Which action is the best engineering choice?
Practice 6 / Práctica 6
Which option would you avoid in a production system?
MODULE 03
🛠️

Environments & Configuration

Development, staging and production should use the same artifact with environment-specific configuration. Separate credentials, endpoints and data while preserving behavior as much as possible.

👁️
See it this way

Environment parity reduces “surprises” between staging and production.

Core ideas

  • Build once
  • Configure per environment
  • Separate secrets
  • Promote tested artifact
Try this
APP_ENV=production
MODEL_VERSION=2026-09
LOG_LEVEL=INFO
DATABASE_URL=...

# same image, different config
✅

Environment parity reduces “surprises” between staging and production.

Practice — 3 cases

Practice 7 / Práctica 7
Which statement best reflects this module?
Practice 8 / Práctica 8
Which action is the best engineering choice?
Practice 9 / Práctica 9
Which option would you avoid in a production system?
MODULE 04
🧪

Networking, Domains & TLS

A production AI API usually sits behind DNS, TLS and a load balancer or managed ingress. Network policy should expose only what clients actually need.

👁️
See it this way

Public by default is not a deployment strategy. Minimize exposed surface.

Core ideas

  • Use HTTPS
  • Keep private services private
  • Route through managed ingress/load balancer
  • Control inbound/outbound traffic
Try this
client -> https://api.example.com
       -> load balancer
       -> AI service
       -> private database/model store
✅

Public by default is not a deployment strategy. Minimize exposed surface.

Practice — 3 cases

Practice 10 / Práctica 10
Which statement best reflects this module?
Practice 11 / Práctica 11
Which action is the best engineering choice?
Practice 12 / Práctica 12
Which option would you avoid in a production system?
MODULE 05
⚙️

CI/CD: Release Without Fear

Continuous integration tests every change; continuous delivery/deployment packages and releases validated changes through an automated pipeline.

👁️
See it this way

Automation reduces manual variance, but only if rollback and verification are designed into the pipeline.

Core ideas

  • Test code
  • Build image
  • Scan/package
  • Deploy and verify
Try this
git push
  -> tests
  -> docker build
  -> push registry
  -> deploy staging
  -> smoke test
  -> promote production
✅

Automation reduces manual variance, but only if rollback and verification are designed into the pipeline.

Practice — 3 cases

Practice 13 / Práctica 13
Which statement best reflects this module?
Practice 14 / Práctica 14
Which action is the best engineering choice?
Practice 15 / Práctica 15
Which option would you avoid in a production system?
MODULE 06
📡

Observability for AI Services

Logs, metrics and traces tell you whether the system is healthy. AI adds extra signals: model latency, token usage, retrieval quality, drift and business outcomes.

👁️
See it this way

If you cannot observe it, you cannot operate it responsibly.

Core ideas

  • Logs explain events
  • Metrics quantify health
  • Traces connect distributed calls
  • AI metrics measure model/application quality
Try this
metrics:
  request_rate
  p95_latency
  error_rate
  cpu_memory
  model_confidence
  token_cost
✅

If you cannot observe it, you cannot operate it responsibly.

Practice — 3 cases

Practice 16 / Práctica 16
Which statement best reflects this module?
Practice 17 / Práctica 17
Which action is the best engineering choice?
Practice 18 / Práctica 18
Which option would you avoid in a production system?
MODULE 07
🛡️

Scaling, Reliability & Resilience

Production services need enough capacity and graceful failure behavior. Scale horizontally when possible, use timeouts/retries carefully and design health checks and rollback paths.

👁️
See it this way

Retries can amplify outages. Reliability patterns must be applied with context.

Core ideas

  • Autoscale based on useful signals
  • Use timeouts
  • Retry only safe operations
  • Design rollback and redundancy
Try this
replicas: 3
autoscale: 3 -> 20
timeout: 5s
health: /health
rollback: previous image tag
✅

Retries can amplify outages. Reliability patterns must be applied with context.

Practice — 3 cases

Practice 19 / Práctica 19
Which statement best reflects this module?
Practice 20 / Práctica 20
Which action is the best engineering choice?
Practice 21 / Práctica 21
Which option would you avoid in a production system?
MODULE 08
🚀

Cost, Governance & Production Readiness

Cloud makes capacity easy to consume, not free. AI workloads can create large compute, GPU and token bills, so cost visibility and governance belong in architecture decisions.

👁️
See it this way

A production system is not finished when it deploys. It is finished when the team can operate, secure, measure and afford it.

Core ideas

  • Tag resources
  • Set budgets/alerts
  • Right-size compute
  • Track cost per useful outcome
Try this
unit_economics = monthly_ai_cost / successful_business_outcomes
print(unit_economics)
✅

A production system is not finished when it deploys. It is finished when the team can operate, secure, measure and afford it.

Practice — 3 cases

Practice 22 / Práctica 22
Which statement best reflects this module?
Practice 23 / Práctica 23
Which action is the best engineering choice?
Practice 24 / Práctica 24
Which option would you avoid in a production system?
5-Question Knowledge Check

Can you explain the core ideas clearly?

Open each item only after answering it in your own words.

1. What is the simplest cloud principle to remember?

Use managed building blocks to run your workload, but keep explicit control of configuration, security, reliability and cost.

2. Why promote the same image across environments?

It preserves artifact consistency and reduces environment-specific rebuild differences.

3. What does CI/CD reduce?

Manual variance in testing, packaging and releasing software.

4. What should AI observability include beyond CPU and errors?

Application/model quality signals such as latency, token/cost usage, retrieval quality, drift and business outcomes.

5. When is a deployment truly production-ready?

When it can be operated, secured, observed, scaled, rolled back and financed sustainably.

Decision Guide

What should you reach for?

NeedRecommended approach
Run simple container APIManaged container platform
Automate releaseCI/CD pipeline
See production healthLogs + metrics + traces
Control cloud spendBudgets + unit economics

Certificate of Participation

Complete at least 12 of the 24 practice cases (50%) and enter your name.

0 / 24 • 0%