Learn the estimator API, preprocessing, train/test splits, pipelines, regression, classification, evaluation, cross-validation and hyperparameter tuning—the practical workflow behind trustworthy machine-learning experiments.
Enter your name. Complete at least 12 practice cases to unlock the Certificate of Participation.
Scikit-learn organizes machine learning around a small, consistent estimator interface. Most models learn with fit(), generate outputs with predict(), and expose parameters through get_params()/set_params().
Treat the estimator API as a workflow contract: prepare data, fit on training data, evaluate on unseen data.
from sklearn.linear_model import LinearRegression
model = LinearRegression()
model.fit(X_train, y_train)
preds = model.predict(X_test)Treat the estimator API as a workflow contract: prepare data, fit on training data, evaluate on unseen data.
A test set estimates how a fitted workflow may behave on unseen data. Any step that learns from the full dataset before the split can leak information and make evaluation look better than reality.
A clean split is not housekeeping—it is part of the scientific design of the experiment.
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, random_state=42, stratify=y
)A clean split is not housekeeping—it is part of the scientific design of the experiment.
Pipelines chain transformations and a final estimator so the same fitted preprocessing is applied consistently during validation and prediction.
If preprocessing is required at prediction time, put it inside the pipeline—not in a forgotten notebook cell.
from sklearn.compose import ColumnTransformer
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler, OneHotEncoder
prep = ColumnTransformer([
("num", StandardScaler(), num_cols),
("cat", OneHotEncoder(handle_unknown="ignore"), cat_cols)
])
pipe = Pipeline([("prep", prep), ("model", LogisticRegression(max_iter=1000))])If preprocessing is required at prediction time, put it inside the pipeline—not in a forgotten notebook cell.
Regression models estimate numeric outcomes such as cost, time, demand or revenue. Start with a simple baseline, then evaluate errors in units the business can interpret.
A regression metric is useful only when you can explain what a typical error means operationally.
from sklearn.metrics import mean_absolute_error, mean_squared_error, r2_score
model = LinearRegression().fit(X_train, y_train)
p = model.predict(X_test)
print(mean_absolute_error(y_test, p))
print(mean_squared_error(y_test, p) ** 0.5)
print(r2_score(y_test, p))A regression metric is useful only when you can explain what a typical error means operationally.
Classification models estimate discrete outcomes such as yes/no, category or risk class. Probabilities often provide more decision value than hard labels alone.
A class probability is not a decision by itself; the decision threshold belongs to the business context.
from sklearn.linear_model import LogisticRegression
clf = LogisticRegression(max_iter=1000)
clf.fit(X_train, y_train)
labels = clf.predict(X_test)
proba = clf.predict_proba(X_test)[:, 1]A class probability is not a decision by itself; the decision threshold belongs to the business context.
Good evaluation uses metrics aligned with the business question and repeated validation when one split may be unstable. Cross-validation estimates performance across multiple train/validation partitions.
Choose metrics before looking at the result whenever possible; otherwise the metric can become part of the overfitting.
from sklearn.model_selection import cross_val_score
from sklearn.metrics import classification_report
scores = cross_val_score(pipe, X, y, cv=5, scoring="f1")
print(scores.mean(), scores.std())
print(classification_report(y_test, pipe.predict(X_test)))Choose metrics before looking at the result whenever possible; otherwise the metric can become part of the overfitting.
Hyperparameters control model behavior before fitting. GridSearchCV and RandomizedSearchCV evaluate candidate settings using cross-validation while keeping the final test set untouched.
Tuning is model selection. If you repeatedly inspect the test set during tuning, it stops being a true test set.
from sklearn.model_selection import RandomizedSearchCV
search = RandomizedSearchCV(
pipe,
{"model__C": [0.01, 0.1, 1, 10, 100]},
n_iter=5, cv=5, scoring="f1", random_state=42
)
search.fit(X_train, y_train)
print(search.best_params_)Tuning is model selection. If you repeatedly inspect the test set during tuning, it stops being a true test set.
A useful model is a reproducible fitted workflow that can receive future rows, apply the same preprocessing and generate predictions under controlled versions and assumptions.
Model files are executable artifacts in practice: load only trusted files and control the environment that produced them.
import joblib
joblib.dump(search.best_estimator_, "model_pipeline.joblib")
loaded = joblib.load("model_pipeline.joblib")
future_pred = loaded.predict(new_rows)Model files are executable artifacts in practice: load only trusted files and control the environment that produced them.
Open each item only after answering it in your own words.
fit() learns from training data; predict() uses the fitted estimator to generate outputs for new feature rows.
So validation and future prediction use the exact same fitted preprocessing sequence without leakage or forgotten steps.
Because tuning is model selection; repeatedly using the test set would bias the final performance estimate.
It estimates how performance varies across multiple train/validation splits rather than relying on one partition.
Controlled data lineage, code, random states, package versions, fitted preprocessing and a documented evaluation procedure.
| Need | Scikit-learn | XGBoost | LightGBM / CatBoost |
|---|---|---|---|
| Transparent linear baseline | Excellent starting point | Often unnecessary | Often unnecessary |
| General preprocessing + many model families | Strong | Usually wrapped through sklearn API | Usually wrapped through sklearn API |
| Tabular boosting at scale | Available via HistGradientBoosting | Strong specialization | Strong specialization |
| Workflow composition and validation | Core strength | Integrates well | Integrates well |
Choose the tool from the problem, data, validation evidence and delivery constraints—not from popularity alone.
The technical concepts follow Scikit-learn’s official user guide and API documentation for estimators, model selection, pipelines, preprocessing, metrics and persistence.
“I can build end-to-end Scikit-learn workflows with reproducible splits, preprocessing pipelines, regression/classification models, appropriate metrics, cross-validation, hyperparameter search and controlled prediction.”
Complete at least 12 of the 24 practice cases (50%) and enter your name.