Python Data Science Library Mastery • Training 03
Article-Training • Data Visualization Foundations

Matplotlib

Turn Data into Visual Evidence: The Foundation of Python Plotting

Learn to choose, build and export charts that answer analytical questions—not just produce graphics.

Question → Choose Visual → Figure/Axes → Encode → Label → Explain → Export
FIG
AX
X
Y
↗
●
▮
▥
↓ Choose → Encode → Explain
Numbers → visual pattern → decision
8learning modules
24interactive practices
5rapid review questions
50%certificate threshold
Learning target
Build enough Matplotlib fluency to choose chart types, explain visual patterns and export reproducible analytical figures.
Practice progress0 / 24

Complete 12 of 24 practices (50%) and enter your name to unlock the Certificate of Participation.

MODULE 01
🧭

The Matplotlib Mental Model: Figure, Axes and Artists

Matplotlib is Python’s foundational plotting library. The most useful mental model is simple: a Figure is the full canvas, an Axes is one plotting area inside it, and the visible elements—lines, text, markers, bars—are Artists placed on that canvas.

👁️
See it this way

Think like a report designer: the Figure is the page, the Axes is the chart panel, and every label, line and marker is an object you can control.

Core ideas

  • Figure = the complete visual canvas
  • Axes = one plotting region with its own x/y coordinate system
  • The object-oriented pattern starts with fig, ax = plt.subplots()
  • pyplot is convenient; Axes methods scale better for controlled work
Try this
import matplotlib.pyplot as plt

fig, ax = plt.subplots()

ax.plot([1, 2, 3], [72, 81, 95])
ax.set_title("Service Resolution Score")
ax.set_xlabel("Week")
ax.set_ylabel("Score")

plt.show()
✅

For quick exploration, pyplot is fine. For reusable business charts, dashboards, reports or multiple panels, prefer explicit Figure/Axes objects.

Practice the decision, not just the syntax

Practice 1
Which object represents one plotting region?
Practice 2
What does plt.subplots() commonly return?
Practice 3
Which pattern is usually easiest to maintain in larger plotting code?
MODULE 02
📈

Line Charts: Show Change Across an Ordered Sequence

A line chart is strongest when the horizontal axis has meaningful order—especially time. It helps the eye follow direction, turning points and rate of change.

👁️
See it this way

If a manager asks “Are response times improving week by week?”, the line is not decoration—it encodes continuity across time.

Core ideas

  • ax.plot(x, y) creates a line plot
  • Markers can reveal individual observations
  • Multiple lines should answer a real comparison question
  • Always label units and make the time grain clear
Try this
import matplotlib.pyplot as plt

weeks = [1, 2, 3, 4, 5]
avg_days = [6.2, 5.8, 5.1, 4.9, 4.4]

fig, ax = plt.subplots()
ax.plot(weeks, avg_days, marker="o")
ax.set_title("Average Resolution Time")
ax.set_xlabel("Week")
ax.set_ylabel("Days")
ax.grid(True, alpha=0.25)

plt.show()
✅

Use a line when connection between adjacent x-values is meaningful. Do not connect unrelated categories just because a line looks smooth.

Practice the decision, not just the syntax

Practice 4
When is a line chart usually the best first choice?
Practice 5
What does marker='o' add to a line plot?
Practice 6
Why can connecting unrelated categories with a line be misleading?
MODULE 03
🔵

Scatter Plots: See Relationships, Clusters and Outliers

Scatter plots place one numeric variable on each axis. They are ideal for asking whether two measures move together, whether groups form, and which observations behave unusually.

👁️
See it this way

Imagine each point is one city service request type: x = volume, y = average days. The upper-right corner immediately exposes high-volume, slow-resolution areas.

Core ideas

  • ax.scatter(x, y) displays paired numeric observations
  • Position is usually more informative than decorative effects
  • Alpha can reduce overplotting when many points overlap
  • Annotate only points that matter to the decision
Try this
import matplotlib.pyplot as plt

volume = [120, 210, 95, 310, 180]
avg_days = [2.1, 3.8, 1.9, 6.2, 4.0]

fig, ax = plt.subplots()
ax.scatter(volume, avg_days, s=70, alpha=0.75)
ax.set_title("Request Volume vs Resolution Time")
ax.set_xlabel("Monthly Requests")
ax.set_ylabel("Average Days")

plt.show()
✅

Scatter does not prove causation. It reveals structure worth investigating: direction, strength, clusters and unusual points.

Practice the decision, not just the syntax

Practice 7
What kind of variables are most natural for a basic scatter plot?
Practice 8
What does alpha help with when points overlap?
Practice 9
A scatter plot shows two variables rising together. What can you conclude immediately?
MODULE 04
📊

Bar Charts: Compare Categories Without Losing the Question

Bar charts compare magnitudes across discrete categories. Their power comes from a shared baseline and easy length comparison, making them excellent for departments, products, regions or status groups.

👁️
See it this way

If the question is “Which department has the most open requests?”, the bar length should make the answer visible before the reader studies exact labels.

Core ideas

  • ax.bar(categories, values) creates vertical bars
  • ax.barh() is useful for long category names
  • Sorting can make ranking patterns easier to scan
  • A zero baseline is normally important for fair length comparison
Try this
import matplotlib.pyplot as plt

departments = ["Police", "Fire", "Parks", "Public Works"]
open_cases = [142, 88, 61, 173]

fig, ax = plt.subplots()
ax.barh(departments, open_cases)
ax.set_title("Open Service Requests")
ax.set_xlabel("Open Requests")

plt.show()
✅

Choose bars for category magnitude. If your x-axis is continuous time, a line may communicate the pattern more naturally.

Practice the decision, not just the syntax

Practice 10
What does a bar chart primarily encode?
Practice 11
When can barh() be especially useful?
Practice 12
Why is sorting bars often helpful?
MODULE 05
📦

Histograms: Understand the Shape of a Distribution

A histogram groups numeric observations into bins and counts how many fall in each interval. It answers a different question from a bar chart: not “which category is larger?” but “how are numeric values distributed?”

👁️
See it this way

Average resolution time can hide operational pain. A histogram can reveal whether most cases close quickly while a long tail of cases remains open for weeks.

Core ideas

  • ax.hist(values, bins=...) groups observations into intervals
  • Bin choice affects how much structure you can see
  • Histograms reveal skew, spread, peaks and possible outliers
  • Compare distributions carefully when sample sizes differ
Try this
import matplotlib.pyplot as plt

days_open = [1, 2, 2, 3, 3, 4, 4, 5, 7, 9, 12, 18, 27]

fig, ax = plt.subplots()
ax.hist(days_open, bins=6, edgecolor="black")
ax.set_title("Distribution of Days Open")
ax.set_xlabel("Days Open")
ax.set_ylabel("Number of Requests")

plt.show()
✅

A histogram summarizes a distribution; it does not preserve the identity of every observation. Use a scatter, strip-style view or table when individual identity matters.

Practice the decision, not just the syntax

Practice 13
What does a histogram group numeric values into?
Practice 14
What can a histogram reveal that a single mean can hide?
Practice 15
What happens when you change the number of bins?
MODULE 06
🏷️

Labels, Legends, Annotations and Scales: Make the Chart Explain Itself

A technically correct chart can still fail if the reader cannot identify the metric, units, comparison or key exception. Labels and annotations should reduce interpretation effort—not decorate the page.

👁️
See it this way

The best annotation answers “why should I look here?” A note such as “system outage” next to a spike is more useful than labeling every point.

Core ideas

  • set_title(), set_xlabel(), set_ylabel() provide context
  • legend() identifies multiple series when labels are supplied
  • annotate() calls attention to a specific point or event
  • set_yscale('log') can help with orders-of-magnitude differences, but must be clearly understood
Try this
import matplotlib.pyplot as plt

months = ["May", "Jun", "Jul", "Aug", "Sep"]
incidents = [42, 47, 51, 93, 55]

fig, ax = plt.subplots()
ax.plot(months, incidents, marker="o", label="Incidents")
ax.set_title("Monthly Incident Volume")
ax.set_ylabel("Incidents")
ax.legend()

ax.annotate(
    "System outage",
    xy=("Aug", 93),
    xytext=("Jun", 105),
    arrowprops={"arrowstyle": "->"}
)

plt.show()
✅

Every extra label has a cost. Add text when it changes understanding; remove it when it only repeats what the chart already shows.

Practice the decision, not just the syntax

Practice 16
Which method adds a title to an Axes?
Practice 17
What is the main purpose of a legend?
Practice 18
When is an annotation most useful?
MODULE 07
🧩

Subplots and Layouts: Build a Visual Story, Not a Chart Pile

Subplots let several related views share one Figure. This is powerful when the panels answer parts of one question—trend, distribution, category comparison—but weak when unrelated charts are crowded together.

👁️
See it this way

A useful operations page might show: top-left = weekly trend, top-right = department ranking, bottom = distribution of resolution days. Each panel adds a different angle to the same operational story.

Core ideas

  • plt.subplots(rows, cols) can create an array of Axes
  • figsize controls the overall Figure dimensions
  • constrained_layout=True helps spacing
  • Shared axes can support fair comparison across panels
Try this
import matplotlib.pyplot as plt

fig, axes = plt.subplots(
    1, 2,
    figsize=(10, 4),
    constrained_layout=True
)

axes[0].plot([1, 2, 3, 4], [8, 7, 6, 5], marker="o")
axes[0].set_title("Avg Days by Week")

axes[1].bar(["A", "B", "C"], [42, 31, 18])
axes[1].set_title("Open Cases by Team")

plt.show()
✅

Use multiple panels when comparison benefits from proximity. If each chart answers a different audience or decision, separate pages may be clearer.

Practice the decision, not just the syntax

Practice 19
What does plt.subplots(1, 2) create conceptually?
Practice 20
What does figsize control?
Practice 21
Why can constrained_layout=True be useful?
MODULE 08
🚀

From Analysis to Deliverable: Export, Reuse and Connect to Pandas

The final skill is not drawing a chart—it is producing a reproducible visual deliverable. A strong workflow prepares data, builds the Figure, validates labels and scales, and exports a file suitable for reports, web pages or presentations.

👁️
See it this way

Think pipeline, not screenshot: raw table → Pandas summary → Matplotlib Figure → validated export. If the data changes next week, the same code should rebuild the chart.

Core ideas

  • fig.savefig() exports PNG, PDF, SVG and other supported formats
  • dpi matters for raster output such as PNG
  • bbox_inches='tight' can reduce unwanted clipping/margins
  • Separate data preparation from chart formatting for easier reuse
Try this
import pandas as pd
import matplotlib.pyplot as plt

df = pd.DataFrame({
    "Department": ["Fire", "Parks", "Police", "Public Works"],
    "Open": [88, 61, 142, 173]
})

summary = df.sort_values("Open")

fig, ax = plt.subplots(figsize=(8, 4))
ax.barh(summary["Department"], summary["Open"])
ax.set_title("Open Requests by Department")
ax.set_xlabel("Open Requests")

fig.savefig(
    "open_requests.png",
    dpi=300,
    bbox_inches="tight"
)

plt.show()
✅

Export the Figure from code instead of relying on screenshots. That preserves repeatability, dimensions and output quality.

Practice the decision, not just the syntax

Practice 22
Which method exports a Figure to a file?
Practice 23
Why is separating data preparation from plotting useful?
Practice 24
What is the strongest production mindset for recurring charts?
5-Question Knowledge Check

Can you explain the visual decision before you write the code?

Open each item only after answering it in your own words. These are review prompts; the 24 interactive practices above drive certificate progress.

1. What is the difference between a Figure and an Axes?

The Figure is the complete canvas. An Axes is one plotting region inside that Figure.

2. When would you choose line, bar or scatter?

Line for ordered progression; bars for categories; scatter for numeric relationships.

3. Why is a histogram not the same as a bar chart?

A histogram groups numeric values into intervals; bars compare discrete categories.

4. What makes an annotation useful?

It explains a meaningful event or exception that changes interpretation.

5. What turns a chart into a reproducible deliverable?

A controlled code pipeline that can regenerate the visual from refreshed data.

Decision Guide

Matplotlib, Seaborn or Plotly?

NeedMatplotlibSeabornPlotly
Precise low-level control of static chartsStrongBuilt on MatplotlibDifferent model
Fast statistical visualization from tidy dataPossibleStrongPossible
Interactive hover, zoom and web-first chartsLimited by defaultLimited by defaultStrong
Foundation for learning Python visualization conceptsExcellentHelpful next layerUseful complementary tool

These tools overlap. Matplotlib is the foundational layer in this series.

Official Sources & Further Learning

Grounded in the official Matplotlib documentation

The technical concepts in this training follow Matplotlib’s official documentation and tutorials.

Market Skills

What you should be able to say after this training

“I can build clear Matplotlib visuals using the Figure/Axes model, choose appropriate chart types, add meaningful labels and annotations, compose multiple panels, and export reproducible figures.”

Certificate of Participation

Complete at least 12 of the 24 practice cases (50%) and enter your name.

0 / 24 • 0%