Learn the core PyTorch workflow: tensors and devices, Dataset/DataLoader, nn.Module, loss and optimizer, autograd, train/eval modes, checkpointing and transfer learning.
Enter your name. Complete at least 12 practice cases to unlock the Certificate of Participation.
PyTorch feels like regular Python while autograd records differentiable tensor operations needed for learning.
Explicit code makes behavior visible, but it also makes you responsible for the training contract.
import torch
x = torch.tensor([[1.,2.],[3.,4.]], requires_grad=True)
print(x.shape)Explicit code makes behavior visible, but it also makes you responsible for the training contract.
Dataset defines how examples are accessed; DataLoader handles batching, shuffling and iteration.
The data loader is part of model performance because it controls what the model sees and how fast it sees it.
from torch.utils.data import TensorDataset, DataLoader
ds = TensorDataset(X, y)
loader = DataLoader(ds, batch_size=32, shuffle=True)The data loader is part of model performance because it controls what the model sees and how fast it sees it.
nn.Module is the standard container for parameters, layers and forward computation.
Keep the forward path readable; complexity should represent the problem, not accidental plumbing.
class Net(torch.nn.Module):
def __init__(self):
super().__init__()
self.fc = torch.nn.Linear(20, 1)
def forward(self, x):
return self.fc(x)Keep the forward path readable; complexity should represent the problem, not accidental plumbing.
Training is an explicit cycle: predict, compute loss, clear old gradients, backpropagate and update parameters.
The order matters; PyTorch gradients accumulate unless you clear them intentionally.
opt.zero_grad()
pred = model(xb)
loss = loss_fn(pred, yb)
loss.backward()
opt.step()The order matters; PyTorch gradients accumulate unless you clear them intentionally.
Autograd constructs a reverse-mode differentiation graph from tensor operations, then applies the chain rule during backward().
Gradient literacy helps debug exploding, missing or unintentionally retained computation.
y = (x ** 2).sum()
y.backward()
print(x.grad)Gradient literacy helps debug exploding, missing or unintentionally retained computation.
Some layers behave differently during training and inference, so model.train() and model.eval() are operationally meaningful.
Evaluation mode is a correctness requirement for layers such as dropout and batch normalization.
model.eval()
with torch.no_grad():
preds = model(X_val)Evaluation mode is a correctness requirement for layers such as dropout and batch normalization.
Move the model and each batch to the same device, then measure the actual bottleneck before optimizing.
Device errors are usually contract errors: tensors participating in one operation must live where the operation expects them.
device = torch.device('cuda' if torch.cuda.is_available() else 'cpu')
model.to(device)Device errors are usually contract errors: tensors participating in one operation must live where the operation expects them.
PyTorch commonly saves state_dict weights, then rebuilds the architecture and loads those states for inference or fine-tuning.
A checkpoint without its architecture, preprocessing and version context is an incomplete artifact.
torch.save(model.state_dict(), 'model.pt')
model.load_state_dict(torch.load('model.pt', weights_only=True))A checkpoint without its architecture, preprocessing and version context is an incomplete artifact.
Open each item only after answering it in your own words.
Because gradients accumulate by default across backward passes.
A standard container that registers parameters/layers and defines forward computation.
It records differentiable operations into a graph used to compute gradients during backward.
To activate evaluation behavior for modules such as dropout and batch normalization.
The architecture/configuration, preprocessing contract, version information and validation evidence.
| Need | Use this training when… | Control |
|---|---|---|
| Fast baseline | You need a defensible first model and workflow | Validate independently |
| Custom behavior | You need to move below the high-level API | Add complexity only for a requirement |
| Production | Production | Version data, code, artifacts and monitoring |
The technical concepts and code patterns in this training follow official project documentation. Validate package versions and environment compatibility before production use.
“I can build PyTorch data loaders and nn.Module models, write and debug explicit training loops, reason about autograd and devices, evaluate correctly and package checkpoints for reproducible deployment or fine-tuning.”
Complete at least 12 of the 24 practice cases (50%) and enter your name.