Learn PyCaret’s experiment workflow: setup, model comparison, tuning, preprocessing, analysis, prediction and model persistence—without losing validation discipline.
Enter your name. Complete at least 12 practice cases to unlock the Certificate of Participation.
PyCaret wraps common machine-learning steps into an experiment. setup() defines the target, split and preprocessing contract before model comparison begins.
Low-code does not mean low-accountability: the split, target, metric and preprocessing choices still belong to you.
from pycaret.classification import *
s = setup(data=df, target='churn', session_id=42)Low-code does not mean low-accountability: the split, target, metric and preprocessing choices still belong to you.
compare_models() trains and ranks candidate estimators with cross-validation so you can establish a defensible baseline quickly.
The top row is not automatically the production winner; metric fit, latency, stability and interpretability still matter.
best = compare_models(sort='F1')
print(best)The top row is not automatically the production winner; metric fit, latency, stability and interpretability still matter.
Use create_model() when you want a specific estimator and tune_model() when you want PyCaret to search its hyperparameters under the experiment rules.
Hyperparameter search can overfit the validation process; keep a final untouched evaluation step.
lr = create_model('lr')
tuned_lr = tune_model(lr, optimize='F1')Hyperparameter search can overfit the validation process; keep a final untouched evaluation step.
setup() can automate imputation, encoding and other transformations. The key is to make those transformations reproducible and leakage-safe.
Automation is valuable only when you can explain what was transformed and reproduce it later.
s = setup(
data=df, target='churn',
normalize=True,
session_id=42
)Automation is valuable only when you can explain what was transformed and reproduce it later.
PyCaret can generate diagnostic plots and evaluation views. Use them to understand errors, not merely to decorate a notebook.
A single aggregate score can hide expensive mistakes in important subgroups.
plot_model(best, plot='confusion_matrix')
plot_model(best, plot='auc')A single aggregate score can hide expensive mistakes in important subgroups.
After model selection, finalize_model() can retrain on the full experiment data, predict_model() scores new rows, and save_model() persists the pipeline.
Delivery is a contract: the incoming schema and preprocessing assumptions must match training.
final_best = finalize_model(best)
preds = predict_model(final_best, data=new_data)
save_model(final_best, 'churn_pipeline')Delivery is a contract: the incoming schema and preprocessing assumptions must match training.
PyCaret supports both functional calls and experiment objects. The object-oriented API is useful when you need multiple isolated experiments in one process.
Choose the API style that makes state and reproducibility clearest for your workflow.
from pycaret.classification import ClassificationExperiment
exp = ClassificationExperiment()
exp.setup(df, target='churn', session_id=42)
best = exp.compare_models()Choose the API style that makes state and reproducibility clearest for your workflow.
PyCaret accelerates experimentation, but production still requires independent validation, monitoring, security, governance and rollback plans.
The goal is not “automate everything”; it is “automate repeatable search while preserving human control.”
# AutoML shortens search; governance closes the loop.
# Validate → document → deploy → monitor → retrainThe goal is not “automate everything”; it is “automate repeatable search while preserving human control.”
Open each item only after answering it in your own words.
The experiment configuration: target, data split, preprocessing and other rules used by later functions.
To build a cross-validated baseline across multiple candidate estimators quickly.
Repeatedly optimizing against the same validation process can adapt to its noise.
The preprocessing pipeline, schema assumptions, package environment and model artifact.
When you need isolated, reusable experiment state—especially for multiple experiments or applications.
| Need | Tool | Focus |
|---|---|---|
| Low-code end-to-end experiment workflow | PyCaret | ★ Current training |
| Scikit-learn pipeline search + ensembles | auto-sklearn | |
| Scalable platform + leaderboard + stacked ensembles | H2O AutoML | |
| Evolutionary pipeline structure search | TPOT | |
| Hyperparameter optimization for your chosen model/code | Optuna | |
| Budget-aware fast AutoML search | FLAML |
Choose from the problem, validation evidence, compute budget and delivery constraints—not from popularity alone.
The technical concepts and code patterns in this training follow the project’s official documentation. Validate package versions and environment compatibility before production use.
“I can structure a PyCaret experiment, compare models with cross-validation, tune candidates, audit preprocessing, analyze model errors, finalize a selected pipeline and persist it for repeatable inference.”
Complete at least 12 of the 24 practice cases (50%) and enter your name.