Move beyond drawing charts one mark at a time. Learn to map tidy data into statistical visuals using relational, distributional, categorical, regression, matrix and faceted views—then finish them with Matplotlib-level control.
Enter your name. Complete at least 12 practice cases to unlock the Certificate of Participation.
Seaborn is a high-level statistical visualization library built on Matplotlib. Its power comes from mapping variables in tidy DataFrames to visual roles such as x, y, hue, size and style.
Think of Seaborn as a visual query language. Instead of manually drawing every mark, you describe which variables should control position, color, size or facets, and Seaborn builds the statistical view.
import seaborn as sns
import pandas as pd
df = pd.DataFrame({
"team": ["A", "A", "B", "B"],
"requests": [120, 135, 95, 110],
"days": [4.2, 3.8, 5.1, 4.6]
})
sns.scatterplot(data=df, x="requests", y="days", hue="team")Use Seaborn when the analytical question involves relationships, distributions, categories or statistical comparisons. Drop to Matplotlib when you need fine-grained layout or artist-level control.
Scatter and line plots answer relational questions. Seaborn adds semantic encodings so one chart can show multiple dimensions without manually splitting the data.
Suppose service volume rises while resolution time also rises. A scatterplot can reveal the relationship, while hue can separate departments and size can encode case severity.
import seaborn as sns
sns.scatterplot(
data=df,
x="requests",
y="days",
hue="team",
size="requests"
)Do not add every semantic channel at once. If hue, size and style all encode different variables, the chart can become harder—not easier—to read.
Averages hide shape. Seaborn distribution plots help you see skew, multiple peaks, tails, density and cumulative behavior.
Two teams can have the same average resolution time but very different operational risk. One may be tightly clustered; the other may have a long tail of delayed cases.
import seaborn as sns
sns.histplot(
data=cases,
x="days_open",
hue="department",
bins=20,
kde=True
)KDE is a smooth estimate, not the raw data itself. Check sample size and the histogram/ECDF before making claims from a smooth curve.
Categorical plots compare numeric outcomes across groups. The right choice depends on whether you want counts, summaries, spread, raw observations or all of them together.
A bar showing average wait time may suggest one team is worse. A boxplot or violin plot may reveal that the real issue is a small group of extreme delays.
import seaborn as sns
sns.boxplot(
data=cases,
x="department",
y="days_open"
)When sample sizes are small, show the observations too. A summary without the underlying points can create a false sense of certainty.
Seaborn often performs statistical aggregation for you. That convenience is powerful, but you must know what estimator and uncertainty are being shown.
A manager asks whether higher workload is associated with slower service. A regression view can show the trend, but the line does not prove that workload causes the delay.
import seaborn as sns
sns.regplot(
data=cases,
x="requests_received",
y="avg_days_open",
scatter_kws={"alpha": 0.6}
)A fitted line is evidence of association under a model, not proof of causation. Bring domain context, diagnostics and experimental design when causal claims matter.
Exploratory work often starts with many variables. Seaborn provides compact views for scanning pairwise relationships, joint distributions and matrix-style patterns.
Before building a model, you may want to know which variables move together, which categories separate naturally, and whether some features are strongly correlated.
import seaborn as sns
num = df[["requests", "days", "satisfaction"]]
corr = num.corr()
sns.heatmap(corr, annot=True, fmt=".2f")Pairplots can become expensive and visually noisy as the number of variables grows. Use them for focused exploration, not for dumping an entire warehouse into one figure.
Faceting repeats the same visual structure across subsets of the data. This reduces clutter and makes comparisons across time periods, regions, teams or categories easier.
Instead of putting twelve service categories into one overloaded chart, create a grid where each panel answers the same question for one category.
import seaborn as sns
sns.relplot(
data=monthly,
x="month",
y="avg_days",
col="department",
col_wrap=3,
kind="line"
)Facets work best when every panel uses the same visual grammar and comparable scales. If each panel needs a different story, separate charts may be clearer.
Professional visualization is not just a chart call. It is a repeatable pipeline: clean data, choose the analytical view, apply a coherent theme, validate labels and scales, then export.
Your chart may go to a dashboard screenshot, article, report or executive deck. The same code should be able to regenerate the visual when the data refreshes.
import seaborn as sns
import matplotlib.pyplot as plt
sns.set_theme(style="whitegrid", context="talk")
ax = sns.barplot(data=summary, x="team", y="avg_days")
ax.set(title="Average Resolution Time", xlabel="Team", ylabel="Days")
plt.tight_layout()
plt.savefig("resolution_time.png", dpi=200, bbox_inches="tight")Style should support interpretation. Accessible labels, readable contrast and accurate scales matter more than decorative effects.
Open each item only after answering it in your own words.
Axes-level functions draw into one Axes; figure-level functions manage a larger figure structure.
Tidy columns map naturally to semantic roles.
Use a boxplot when distribution shape and spread matter.
Do not infer causation from the fitted line alone.
Seaborn handles the statistical view; Matplotlib provides precise finishing control.
| Need | Matplotlib | Seaborn | Plotly |
|---|---|---|---|
| Precise low-level static control | Strong | Uses Matplotlib underneath | Different model |
| Fast statistical visualization from tidy data | Possible | Strong | Possible |
| Small multiples / faceting | Manual or custom | Strong | Strong |
| Interactive hover and browser-first charts | Limited by default | Limited by default | Strong |
The tools overlap; choose based on the analytical and delivery need.
The technical concepts in this training follow Seaborn’s official documentation.
“I can use Seaborn with tidy pandas data to build and interpret statistical visualizations, then finish them reproducibly with Matplotlib.”
Complete at least 12 of the 24 practice cases (50%) and enter your name.