Good place to stop if you're short on time — step 3 picks up here.
Why might a lead inspect a sample instead of every row?
To audit quality efficiently when a full review is too costly or slow.
What is a gold label?
A trusted reference decision used to check evaluator or system output.
Is percentage agreement enough when one label dominates?
Not always; high agreement can occur by chance when most items receive the same label.
Does high observed agreement prove accuracy, and what does kappa adjust for?
No. Evaluators can agree and both be wrong. Kappa adjusts for expected chance agreement from their marginal label shares; it is undefined when expected agreement is one.
Record which ideas needed assistance and revisit their lessons. Self-review is not a scored correct answer.
Trace and rebuild reference agreement, repair a classroom review flag, then independently return evaluator evidence with explicit exclusions, denominators and undefined metrics.
If the rebuild is hard, that is information — not failure.
Lesson idea Tap to fold
What should happen after an agreement rate falls below target?
Show the answer
Inspect disagreement examples, clarify the rubric, calibrate, and measure again.
Keep evidence before judging a metric
Rebuild the reference-agreement table using a boolean correct column, then its mean within each evaluator. A strict rate < 0.75 is only a classroom review flag: exactly 0.75 does not trigger it. A rounded display and a small sample do not justify automatically judging or punishing an evaluator. Inspect disagreements and clarify the rubric first.
The independent task combines validation, first-valid retention and confusion counts. Do not globally deduplicate item_id: two different evaluators rating the same item are two observations. A tuple (evaluator, item_id) is an immutable composite key. A set starts with set(); key in seen tests membership and seen.add(key) records a retained pair. Validate and strip string IDs before reserving a key.
Create fresh groups inside each function call, accumulate TP/FP/FN/TN, then compute each denominator. Return data, not just its printed representation. Keep the input unchanged; report raw, rejected, duplicate and retained counts plus disagreement IDs. Return None for undefined rates or kappa, and round defined metrics to four decimals after calculation. Allow about 59 minutes for four tasks; split before the independent report if needed. The quiz is separate.
# Retain pairs, not globally unique item IDs.
pairs = [("Ava", "A"), ("Ava", "A"), ("Ben", "A")]
seen = set()
kept = []
for key in pairs:
if key in seen:
continue
seen.add(key)
kept.append(key)
print(kept)
print("Same item, separate evaluators:", len(kept))
The duplicate Ava/A pair is removed, while Ben/A is retained. Identity depends on the stated composite key, not just item ID.
Explain Ava’s false-positive row and Ben’s undefined kappa in the independent sample. A passing software check is practical evidence, not a claim of general evaluator quality.
Predict, explain, then test
Answer these three self-reviewed checks before opening the comparisons. They are not scored. Coding checks use unfamiliar inputs under each task’s stated assumptions.
Choose identity
Ava and Ben rate the same item ID. Should one row be dropped as a duplicate?
Compare your answer · self-reviewed
No. Different evaluator/item pairs are separate observations; only later valid occurrences of the same pair are duplicates.
Inspect the boundary
Does exactly 0.75 trigger the classroom below-75% flag?
Compare your answer · self-reviewed
No. The comparison is strictly less than; use the full rate, not the rounded display.
Preserve evidence
Why return None and disagreement IDs alongside rates?
Compare your answer · self-reviewed
None distinguishes absence from zero; disagreement IDs let a person inspect retained evidence and the rubric before interpreting a metric.
Independent transfer: an evaluator evidence report
Write summarize_evaluators(rows) independently from this explicit report policy:
- Write summarize_evaluators(rows). Accept a list, otherwise raise TypeError. Do not change its rows or dictionaries; start fresh on every call.
- Reject a row unless it is a dictionary with nonempty string item_id and evaluator after stripping outer whitespace, and exact string gold/label values pass or fail. Count rejected rows separately; do not convert numbers or booleans into labels. Extra fields are allowed.
- After validation, keep the first valid (evaluator, item_id) pair. A duplicate for that evaluator counts as duplicates; the same item rated by a different evaluator remains separate. Invalid earlier rows do not reserve a pair.
- Return raw_rows, rejected, duplicates, kept and evaluators. evaluators maps each cleaned evaluator name to rows, tp, fp, fn, tn, agreement, precision, recall, kappa and disagreements. Counts are integers; disagreements is an input-order list of retained item IDs where label differs from gold.
- Treat pass as the positive class. TP is gold pass / label pass; FP is gold fail / label pass; FN is gold pass / label fail; TN is gold fail / label fail. Agreement is (tp + tn) / rows, precision is tp / (tp + fp), and recall is tp / (tp + fn).
- For kappa, observed is agreement. Gold-pass share is (tp + fn) / rows; predicted-pass share is (tp + fp) / rows. expected is their product plus the product of their complements. kappa is (observed - expected) / (1 - expected).
- Return None for precision or recall with a zero denominator, and for kappa when expected is one. Round defined metrics to four decimals only after calculation. Empty input returns zero counts and an empty evaluators dictionary. Explain the retained sample and disagreement examples; these rates do not prove general evaluator quality or justify automatic punishment.
Select the independent task in the exercise controls. The quiz’s practical link selects it and opens this brief without running code or awarding evidence. Review validation, precision and recall, and chance-corrected agreement.
Read the unfinished starter without running Python
# Summarize retained evidence, not a leaderboard.
def summarize_evaluators(rows):
# Implement the policy; this placeholder must fail.
return {"kept": 0, "evaluators": {}}
# Keep repeated and cross-evaluator IDs distinct.
rows = [
{"item_id": "a1", "evaluator": "Ava",
"gold": "pass", "label": "pass"},
{"item_id": "a2", "evaluator": "Ava",
"gold": "fail", "label": "pass"},
{"item_id": "a1", "evaluator": "Ava",
"gold": "pass", "label": "fail"},
{"item_id": "a1", "evaluator": "Ben",
"gold": "pass", "label": "pass"},
{"item_id": "", "evaluator": "Ava",
"gold": "fail", "label": "fail"},
]
result = summarize_evaluators(rows)
print("Kept:", result["kept"])
print("Evaluators:", len(result["evaluators"]))
The completed sample prints Kept: 3 and Evaluators: 2 on separate lines. Ava retains two rows: TP 1, FP 1, FN 0, TN 0, agreement 0.5, precision 0.5, recall 1, kappa 0 and disagreements [a2]. Ben retains one TP row: agreement, precision and recall 1; kappa None; no disagreements. Explain why Ben’s perfect observed match does not define kappa. Checks inspect returned data on new, empty, malformed and duplicate inputs, repeated calls and unchanged rows. Your quiz score and practical evidence stay separate.
Choose an exercise to load its prompt.
Loading exercises…
Noted on your route map, with the step you were on. Nothing is lost by parking it — the next review day will bring this idea back, and the flag tells the course where to slow down.
Section 7 checkpoint
Answer five questions, review explanations and use the linked refreshers. The independent evaluator report is a separate practical task; neither requires tutor login.
The tutor mounts here when JavaScript is available. The lesson above stays readable without it.