Python Data Science Library Mastery • Training 06
Article-Training • Machine Learning Foundations

Scikit-learn

Build Reliable Machine Learning Workflows with a Consistent API

Learn the estimator API, preprocessing, train/test splits, pipelines, regression, classification, evaluation, cross-validation and hyperparameter tuning—the practical workflow behind trustworthy machine-learning experiments.

Question → Features/Target → Split → Preprocess → Fit → Evaluate → Tune → Predict
X + y • train/test • pipeline
↓
🧠
✂️
🧹
📈
🎯
📏
🔁
🚀
↓
fit() • predict() • score() • validate
8modules
24interactive practices
50%certificate unlock
6market-ready skills
Your Learning Record

Make the practice count

Enter your name. Complete at least 12 practice cases to unlock the Certificate of Participation.

Practice progress0 / 24
MODULE 01
🧠

The Scikit-learn Mental Model: Estimators, fit() and predict()

Scikit-learn organizes machine learning around a small, consistent estimator interface. Most models learn with fit(), generate outputs with predict(), and expose parameters through get_params()/set_params().

👁️
See it this way

Treat the estimator API as a workflow contract: prepare data, fit on training data, evaluate on unseen data.

Core ideas

  • Estimator objects hold configuration and learned state
  • fit(X, y) learns from training data
  • predict(X) produces outputs for new rows
  • Consistent APIs make models easier to compare
Try this
from sklearn.linear_model import LinearRegression

model = LinearRegression()
model.fit(X_train, y_train)
preds = model.predict(X_test)
✅

Treat the estimator API as a workflow contract: prepare data, fit on training data, evaluate on unseen data.

Practice the decision, not just the syntax

Practice 1
Which method learns model parameters from training data?
Practice 2
What does predict() use?
Practice 3
Why is a common API useful?
MODULE 02
✂️

Train/Test Split and Leakage: Protect the Evaluation

A test set estimates how a fitted workflow may behave on unseen data. Any step that learns from the full dataset before the split can leak information and make evaluation look better than reality.

👁️
See it this way

A clean split is not housekeeping—it is part of the scientific design of the experiment.

Core ideas

  • train_test_split() creates independent partitions
  • Use stratify=y when class proportions should be preserved
  • Set random_state for reproducibility
  • Fit preprocessing only on training data
Try this
from sklearn.model_selection import train_test_split

X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, random_state=42, stratify=y
)
✅

A clean split is not housekeeping—it is part of the scientific design of the experiment.

Practice the decision, not just the syntax

Practice 4
What is the main purpose of a test set?
Practice 5
What can cause data leakage?
Practice 6
When is stratify=y helpful?
MODULE 03
🧹

Preprocessing and Pipelines: Make the Workflow Reproducible

Pipelines chain transformations and a final estimator so the same fitted preprocessing is applied consistently during validation and prediction.

👁️
See it this way

If preprocessing is required at prediction time, put it inside the pipeline—not in a forgotten notebook cell.

Core ideas

  • StandardScaler standardizes numeric features
  • OneHotEncoder handles categorical levels
  • ColumnTransformer applies different steps by column
  • Pipeline prevents train/predict drift
Try this
from sklearn.compose import ColumnTransformer
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler, OneHotEncoder

prep = ColumnTransformer([
    ("num", StandardScaler(), num_cols),
    ("cat", OneHotEncoder(handle_unknown="ignore"), cat_cols)
])
pipe = Pipeline([("prep", prep), ("model", LogisticRegression(max_iter=1000))])
✅

If preprocessing is required at prediction time, put it inside the pipeline—not in a forgotten notebook cell.

Practice the decision, not just the syntax

Practice 7
What does Pipeline help guarantee?
Practice 8
What is ColumnTransformer for?
Practice 9
Why use handle_unknown="ignore" in OneHotEncoder?
MODULE 04
📈

Regression: Predict Continuous Outcomes

Regression models estimate numeric outcomes such as cost, time, demand or revenue. Start with a simple baseline, then evaluate errors in units the business can interpret.

👁️
See it this way

A regression metric is useful only when you can explain what a typical error means operationally.

Core ideas

  • LinearRegression is a transparent baseline
  • MAE reports average absolute error
  • RMSE penalizes larger errors more strongly
  • R² describes variance explained relative to a baseline
Try this
from sklearn.metrics import mean_absolute_error, mean_squared_error, r2_score

model = LinearRegression().fit(X_train, y_train)
p = model.predict(X_test)
print(mean_absolute_error(y_test, p))
print(mean_squared_error(y_test, p) ** 0.5)
print(r2_score(y_test, p))
✅

A regression metric is useful only when you can explain what a typical error means operationally.

Practice the decision, not just the syntax

Practice 10
What kind of target does regression predict?
Practice 11
Which metric is easy to explain in target units?
Practice 12
What is a good first model strategy?
MODULE 05
🎯

Classification: Predict Classes and Probabilities

Classification models estimate discrete outcomes such as yes/no, category or risk class. Probabilities often provide more decision value than hard labels alone.

👁️
See it this way

A class probability is not a decision by itself; the decision threshold belongs to the business context.

Core ideas

  • LogisticRegression is a strong classification baseline
  • predict() returns class labels
  • predict_proba() returns estimated probabilities when supported
  • Decision thresholds should reflect business costs
Try this
from sklearn.linear_model import LogisticRegression

clf = LogisticRegression(max_iter=1000)
clf.fit(X_train, y_train)
labels = clf.predict(X_test)
proba = clf.predict_proba(X_test)[:, 1]
✅

A class probability is not a decision by itself; the decision threshold belongs to the business context.

Practice the decision, not just the syntax

Practice 13
What does predict_proba() provide when supported?
Practice 14
Why might 0.5 not be the best threshold?
Practice 15
What is a common transparent classification baseline?
MODULE 06
📏

Evaluation and Cross-Validation: Measure What Matters

Good evaluation uses metrics aligned with the business question and repeated validation when one split may be unstable. Cross-validation estimates performance across multiple train/validation partitions.

👁️
See it this way

Choose metrics before looking at the result whenever possible; otherwise the metric can become part of the overfitting.

Core ideas

  • Accuracy can hide minority-class problems
  • Precision focuses on predicted positives
  • Recall focuses on captured positives
  • cross_val_score repeats evaluation across folds
Try this
from sklearn.model_selection import cross_val_score
from sklearn.metrics import classification_report

scores = cross_val_score(pipe, X, y, cv=5, scoring="f1")
print(scores.mean(), scores.std())
print(classification_report(y_test, pipe.predict(X_test)))
✅

Choose metrics before looking at the result whenever possible; otherwise the metric can become part of the overfitting.

Practice the decision, not just the syntax

Practice 16
What does cross-validation help estimate?
Practice 17
Which metric emphasizes finding as many true positives as possible?
Practice 18
Why can accuracy mislead on imbalanced data?
MODULE 07
🔁

Hyperparameter Tuning: Search without Cheating

Hyperparameters control model behavior before fitting. GridSearchCV and RandomizedSearchCV evaluate candidate settings using cross-validation while keeping the final test set untouched.

👁️
See it this way

Tuning is model selection. If you repeatedly inspect the test set during tuning, it stops being a true test set.

Core ideas

  • GridSearchCV checks an explicit grid
  • RandomizedSearchCV samples parameter combinations
  • Tune the full pipeline, not isolated steps
  • Keep a final test set for one unbiased evaluation
Try this
from sklearn.model_selection import RandomizedSearchCV

search = RandomizedSearchCV(
    pipe,
    {"model__C": [0.01, 0.1, 1, 10, 100]},
    n_iter=5, cv=5, scoring="f1", random_state=42
)
search.fit(X_train, y_train)
print(search.best_params_)
✅

Tuning is model selection. If you repeatedly inspect the test set during tuning, it stops being a true test set.

Practice the decision, not just the syntax

Practice 19
What should stay untouched during tuning?
Practice 20
What does RandomizedSearchCV do?
Practice 21
How are nested pipeline parameters referenced?
MODULE 08
🚀

Inference, Persistence and Reproducibility

A useful model is a reproducible fitted workflow that can receive future rows, apply the same preprocessing and generate predictions under controlled versions and assumptions.

👁️
See it this way

Model files are executable artifacts in practice: load only trusted files and control the environment that produced them.

Core ideas

  • Persist the complete pipeline when appropriate
  • Record package versions and training data lineage
  • Validate input schema before prediction
  • Monitor performance after deployment
Try this
import joblib

joblib.dump(search.best_estimator_, "model_pipeline.joblib")
loaded = joblib.load("model_pipeline.joblib")
future_pred = loaded.predict(new_rows)
✅

Model files are executable artifacts in practice: load only trusted files and control the environment that produced them.

Practice the decision, not just the syntax

Practice 22
What should usually be persisted when preprocessing is required?
Practice 23
Why record library versions?
Practice 24
What should happen before scoring future rows?
5-Question Knowledge Check

Can you explain the model decision before you write the code?

Open each item only after answering it in your own words.

1. What is the difference between fit() and predict()?

fit() learns from training data; predict() uses the fitted estimator to generate outputs for new feature rows.

2. Why put preprocessing inside a Pipeline?

So validation and future prediction use the exact same fitted preprocessing sequence without leakage or forgotten steps.

3. Why keep a final test set untouched during tuning?

Because tuning is model selection; repeatedly using the test set would bias the final performance estimate.

4. What is cross-validation for?

It estimates how performance varies across multiple train/validation splits rather than relying on one partition.

5. What makes a model workflow reproducible?

Controlled data lineage, code, random states, package versions, fitted preprocessing and a documented evaluation procedure.

Decision Guide

Baseline, Scikit-learn or Specialized Boosting?

NeedScikit-learnXGBoostLightGBM / CatBoost
Transparent linear baselineExcellent starting pointOften unnecessaryOften unnecessary
General preprocessing + many model familiesStrongUsually wrapped through sklearn APIUsually wrapped through sklearn API
Tabular boosting at scaleAvailable via HistGradientBoostingStrong specializationStrong specialization
Workflow composition and validationCore strengthIntegrates wellIntegrates well

Choose the tool from the problem, data, validation evidence and delivery constraints—not from popularity alone.

Official Sources & Further Learning

Grounded in the official Scikit-learn documentation

The technical concepts follow Scikit-learn’s official user guide and API documentation for estimators, model selection, pipelines, preprocessing, metrics and persistence.

Market Skills

What you should be able to say after this training

“I can build end-to-end Scikit-learn workflows with reproducible splits, preprocessing pipelines, regression/classification models, appropriate metrics, cross-validation, hyperparameter search and controlled prediction.”

Certificate of Participation

Complete at least 12 of the 24 practice cases (50%) and enter your name.

0 / 24 • 0%