Python Data Science Library Mastery • Training 04
Article-Training • Statistical Visualization

Seaborn

See Relationships, Distributions and Groups with Statistical Context

Move beyond drawing charts one mark at a time. Learn to map tidy data into statistical visuals using relational, distributional, categorical, regression, matrix and faceted views—then finish them with Matplotlib-level control.

DataFrame → Semantic Mapping → Statistical View → Comparison → Interpretation → Decision
x = workload • y = resolution time • hue = team
↓
x
y
hue
col
📈
📦
🔥
🪟
↓
One dataset → many analytical views
8modules
24interactive practices
50%certificate unlock
6market-ready visual skills
Your Learning Record

Make the practice count

Enter your name. Complete at least 12 practice cases to unlock the Certificate of Participation.

Practice progress0 / 24
MODULE 01
🧭

The Seaborn Mental Model: Statistical Views from Tidy Data

Seaborn is a high-level statistical visualization library built on Matplotlib. Its power comes from mapping variables in tidy DataFrames to visual roles such as x, y, hue, size and style.

👁️
See it this way

Think of Seaborn as a visual query language. Instead of manually drawing every mark, you describe which variables should control position, color, size or facets, and Seaborn builds the statistical view.

Core ideas

  • Works naturally with pandas DataFrames and tidy/long-form data
  • Axes-level functions draw into one Matplotlib Axes
  • Figure-level functions can create grids, legends and facets for you
  • Semantic mappings include x, y, hue, size, style, row and col
Try this
import seaborn as sns
import pandas as pd

df = pd.DataFrame({
    "team": ["A", "A", "B", "B"],
    "requests": [120, 135, 95, 110],
    "days": [4.2, 3.8, 5.1, 4.6]
})

sns.scatterplot(data=df, x="requests", y="days", hue="team")
✅

Use Seaborn when the analytical question involves relationships, distributions, categories or statistical comparisons. Drop to Matplotlib when you need fine-grained layout or artist-level control.

Practice the decision, not just the syntax

Practice 1
What is Seaborn primarily designed for?
Practice 2
Which data shape works especially well with Seaborn?
Practice 3
What does hue usually represent?
MODULE 02
🔗

Relational Plots: Show How Variables Move Together

Scatter and line plots answer relational questions. Seaborn adds semantic encodings so one chart can show multiple dimensions without manually splitting the data.

👁️
See it this way

Suppose service volume rises while resolution time also rises. A scatterplot can reveal the relationship, while hue can separate departments and size can encode case severity.

Core ideas

  • sns.scatterplot() explores relationships between numeric variables
  • sns.lineplot() is useful for ordered data such as time
  • hue, style and size add semantic dimensions
  • relplot() is the figure-level relational interface
Try this
import seaborn as sns

sns.scatterplot(
    data=df,
    x="requests",
    y="days",
    hue="team",
    size="requests"
)
✅

Do not add every semantic channel at once. If hue, size and style all encode different variables, the chart can become harder—not easier—to read.

Practice the decision, not just the syntax

Practice 4
Which function is a natural choice for two numeric variables?
Practice 5
What can size add to a scatterplot?
Practice 6
When is lineplot() especially useful?
MODULE 03
📈

Distributions: See Shape, Spread and Uncertainty

Averages hide shape. Seaborn distribution plots help you see skew, multiple peaks, tails, density and cumulative behavior.

👁️
See it this way

Two teams can have the same average resolution time but very different operational risk. One may be tightly clustered; the other may have a long tail of delayed cases.

Core ideas

  • histplot() shows counts or densities across bins
  • kdeplot() estimates a smooth density curve
  • ecdfplot() shows the cumulative proportion at or below each value
  • displot() can facet distributions across categories
Try this
import seaborn as sns

sns.histplot(
    data=cases,
    x="days_open",
    hue="department",
    bins=20,
    kde=True
)
✅

KDE is a smooth estimate, not the raw data itself. Check sample size and the histogram/ECDF before making claims from a smooth curve.

Practice the decision, not just the syntax

Practice 7
What does histplot() primarily show?
Practice 8
What is KDE?
Practice 9
What does an ECDF show?
MODULE 04
📦

Categorical Views: Compare Groups Without Losing the Distribution

Categorical plots compare numeric outcomes across groups. The right choice depends on whether you want counts, summaries, spread, raw observations or all of them together.

👁️
See it this way

A bar showing average wait time may suggest one team is worse. A boxplot or violin plot may reveal that the real issue is a small group of extreme delays.

Core ideas

  • countplot() counts categorical observations
  • barplot() displays an estimator with uncertainty
  • boxplot() summarizes quartiles and potential outliers
  • violinplot() and stripplot() reveal distribution shape and observations
Try this
import seaborn as sns

sns.boxplot(
    data=cases,
    x="department",
    y="days_open"
)
✅

When sample sizes are small, show the observations too. A summary without the underlying points can create a false sense of certainty.

Practice the decision, not just the syntax

Practice 10
Which plot best summarizes quartiles across categories?
Practice 11
Which plot shows category counts?
Practice 12
Why combine a stripplot with a summary plot?
MODULE 05
🎯

Statistical Estimation: Summaries, Error Bars and Regression

Seaborn often performs statistical aggregation for you. That convenience is powerful, but you must know what estimator and uncertainty are being shown.

👁️
See it this way

A manager asks whether higher workload is associated with slower service. A regression view can show the trend, but the line does not prove that workload causes the delay.

Core ideas

  • barplot() can estimate means or another chosen statistic
  • errorbar controls how uncertainty is represented
  • regplot() adds a fitted regression relationship to one Axes
  • lmplot() is a figure-level interface that can facet regression views
Try this
import seaborn as sns

sns.regplot(
    data=cases,
    x="requests_received",
    y="avg_days_open",
    scatter_kws={"alpha": 0.6}
)
✅

A fitted line is evidence of association under a model, not proof of causation. Bring domain context, diagnostics and experimental design when causal claims matter.

Practice the decision, not just the syntax

Practice 13
What does an error bar communicate?
Practice 14
What does a regression line prove by itself?
Practice 15
Which function adds a regression fit to one Axes?
MODULE 06
🧩

Multivariate Exploration: Pairplots, Jointplots and Heatmaps

Exploratory work often starts with many variables. Seaborn provides compact views for scanning pairwise relationships, joint distributions and matrix-style patterns.

👁️
See it this way

Before building a model, you may want to know which variables move together, which categories separate naturally, and whether some features are strongly correlated.

Core ideas

  • pairplot() scans pairwise relationships across several variables
  • jointplot() combines a bivariate view with marginal distributions
  • heatmap() visualizes matrices such as correlations or confusion matrices
  • Correlation is useful for screening, but it is not a causal explanation
Try this
import seaborn as sns

num = df[["requests", "days", "satisfaction"]]
corr = num.corr()

sns.heatmap(corr, annot=True, fmt=".2f")
✅

Pairplots can become expensive and visually noisy as the number of variables grows. Use them for focused exploration, not for dumping an entire warehouse into one figure.

Practice the decision, not just the syntax

Practice 16
What is pairplot() useful for?
Practice 17
What does heatmap() naturally visualize?
Practice 18
What should you remember about correlation?
MODULE 07
🪟

Faceting: Small Multiples for Better Comparisons

Faceting repeats the same visual structure across subsets of the data. This reduces clutter and makes comparisons across time periods, regions, teams or categories easier.

👁️
See it this way

Instead of putting twelve service categories into one overloaded chart, create a grid where each panel answers the same question for one category.

Core ideas

  • relplot(), catplot() and displot() support row/col faceting
  • FacetGrid is the flexible lower-level grid interface
  • Shared axes help comparisons when scales are comparable
  • Too many facets can create tiny unreadable panels
Try this
import seaborn as sns

sns.relplot(
    data=monthly,
    x="month",
    y="avg_days",
    col="department",
    col_wrap=3,
    kind="line"
)
✅

Facets work best when every panel uses the same visual grammar and comparable scales. If each panel needs a different story, separate charts may be clearer.

Practice the decision, not just the syntax

Practice 19
What is faceting?
Practice 20
Which arguments commonly create facets?
Practice 21
What is a risk of too many facets?
MODULE 08
🎨

Design, Themes and a Reproducible Visualization Workflow

Professional visualization is not just a chart call. It is a repeatable pipeline: clean data, choose the analytical view, apply a coherent theme, validate labels and scales, then export.

👁️
See it this way

Your chart may go to a dashboard screenshot, article, report or executive deck. The same code should be able to regenerate the visual when the data refreshes.

Core ideas

  • sns.set_theme() establishes coherent defaults
  • set_context() adjusts scale for notebook, paper, talk or poster contexts
  • color_palette() and palette arguments control categorical or sequential color mapping
  • Use Matplotlib methods for precise titles, annotations, layout and savefig() export
Try this
import seaborn as sns
import matplotlib.pyplot as plt

sns.set_theme(style="whitegrid", context="talk")

ax = sns.barplot(data=summary, x="team", y="avg_days")
ax.set(title="Average Resolution Time", xlabel="Team", ylabel="Days")
plt.tight_layout()
plt.savefig("resolution_time.png", dpi=200, bbox_inches="tight")
✅

Style should support interpretation. Accessible labels, readable contrast and accurate scales matter more than decorative effects.

Practice the decision, not just the syntax

Practice 22
What does sns.set_theme() help control?
Practice 23
When should you use Matplotlib together with Seaborn?
Practice 24
What makes a visualization workflow reproducible?
5-Question Knowledge Check

Can you explain the visual decision before you write the code?

Open each item only after answering it in your own words.

1. What is the difference between an axes-level and a figure-level Seaborn function?

Axes-level functions draw into one Axes; figure-level functions manage a larger figure structure.

2. Why is tidy data important in Seaborn?

Tidy columns map naturally to semantic roles.

3. When would you choose a boxplot over a barplot?

Use a boxplot when distribution shape and spread matter.

4. What should you avoid claiming from a regression line alone?

Do not infer causation from the fitted line alone.

5. Why combine Seaborn with Matplotlib?

Seaborn handles the statistical view; Matplotlib provides precise finishing control.

Decision Guide

Matplotlib, Seaborn or Plotly?

NeedMatplotlibSeabornPlotly
Precise low-level static controlStrongUses Matplotlib underneathDifferent model
Fast statistical visualization from tidy dataPossibleStrongPossible
Small multiples / facetingManual or customStrongStrong
Interactive hover and browser-first chartsLimited by defaultLimited by defaultStrong

The tools overlap; choose based on the analytical and delivery need.

Official Sources & Further Learning

Grounded in the official Seaborn documentation

The technical concepts in this training follow Seaborn’s official documentation.

Market Skills

What you should be able to say after this training

“I can use Seaborn with tidy pandas data to build and interpret statistical visualizations, then finish them reproducibly with Matplotlib.”

Certificate of Participation

Complete at least 12 of the 24 practice cases (50%) and enter your name.

0 / 24 • 0%