Move from 'this answer looks good' to repeatable evaluation. Build datasets, rubrics, deterministic checks, model-based judges, human review, regression tests and release gates.
Complete 12 of 24 practices (50%) and enter your name to unlock the Certificate of Participation.
An eval begins with an explicit definition of what good behavior means for the task.
If success is vague, scores become arbitrary.
criteria = {'correct': True, 'grounded': True, 'format_ok': True}A score is useful only when it maps to a defined product requirement.
A useful eval set reflects common requests, edge cases, no-answer cases and known risks.
Testing only happy paths creates false confidence.
eval_set = [common_cases, edge_cases, failure_cases]The eval set is a product asset and should evolve with the system.
Some requirements can be measured exactly: schema validity, length, allowed values, citations or required fields.
Use code when the rule is objective instead of asking a judge model to guess.
assert jsonschema.validate(output, schema)
assert latency_ms < 2500Prefer deterministic evaluation whenever the success condition is deterministic.
Model-based judges can score nuanced outputs such as relevance, completeness or style when a deterministic rule is insufficient.
A judge is another model, not an oracle.
score = judge(output, rubric='0-4 relevance')Judge prompts and rubrics need their own evaluation and versioning.
Human review remains important for ambiguous, high-impact or preference-heavy behavior.
Humans are expensive but often define the reference standard for nuanced quality.
review = {'score': 3, 'reason': 'supported and concise'}Human evaluation is strongest when the rubric is explicit and reviewers are calibrated.
Every prompt, model or retrieval change can improve one case and break another. Compare versions on the same eval set.
A single better example does not prove the new version is better overall.
compare(v1, v2, eval_set)
assert regressions < thresholdTreat evals as automated regression tests for probabilistic systems.
Complex AI systems can fail at retrieval, planning, tool use or generation. Layered evals identify the real failure point.
Diagnosing the wrong layer wastes tuning effort.
scores = {'retrieval': .9, 'answer': .7, 'task': .6}Use component metrics plus an end-to-end success metric.
Offline evals should influence deployment decisions and continue as online monitoring signals after release.
Evaluation becomes operational when it can block a bad release and detect drift later.
if quality < gate: block_release()
monitor(live_metrics)Connect eval results to release gates and production observability.
Open each item after answering it in your own words.
Define explicit, testable success criteria before choosing metrics.
Use representative cases that cover normal, edge and failure behavior.
Use exact programmatic checks for objective requirements.
Use calibrated rubrics and validate judge agreement with humans.
Use trained reviewers and explicit rubrics for nuanced or high-impact cases.
| Layer | Purpose |
|---|---|
| Define Success Before Testing | Define explicit, testable success criteria before choosing metrics. |
| Build Representative Eval Sets | Use representative cases that cover normal, edge and failure behavior. |
| Deterministic Checks | Use exact programmatic checks for objective requirements. |
| LLM-as-Judge | Use calibrated rubrics and validate judge agreement with humans. |
| Human Evaluation | Use trained reviewers and explicit rubrics for nuanced or high-impact cases. |
| Regression Testing and Comparisons | Compare versions on fixed representative cases and track regressions. |
| Evaluate RAG and Agents by Layer | Evaluate components separately and also measure end-to-end task success. |
| Release Gates and Production Monitoring | Use eval thresholds as release gates and monitor post-release drift. |
Complete at least 12 of the 24 practice cases (50%) and enter your name.
Evaluation metrics and judge models are imperfect. Calibrate them against human judgment, representative tasks and real failure modes.