Good place to stop if you're short on time — step 3 picks up here.
Lesson idea Tap to fold
Is percentage agreement enough when one label dominates?
Show the answer
Not always; high agreement can occur by chance when most items receive the same label.
Agreement is not automatically accuracy
Observed agreement is the matching share of aligned labels, including both pass-pass and fail-fail. Two evaluators may agree and both be wrong. Agreement with a chosen reference answers another question. The introductory pair tasks assume nonempty, equal-length aligned pass/fail lists.
Cohen’s kappa adjusts observed agreement using the marginal label shares. If a_pass and b_pass are their pass proportions, expected agreement is a_pass * b_pass + (1 - a_pass) * (1 - b_pass). Include both matching classes. Kappa is (observed - expected) / (1 - expected). Negative values indicate agreement below this marginal chance model, not a broken calculation.
When expected is one, kappa is undefined even if observed agreement is perfect: the denominator is zero. The introductory helper assumes expected < 1; this worked helper and the independent report return None for the undefined case. The chance model is a convention, not proof that people actually guessed independently. Sample composition and the reference matter; do not use one universal cutoff to judge people.
# Guard the chance-correction denominator.
def kappa(observed, expected):
if expected == 1:
return None
return (observed - expected) / (1 - expected)
a_pass, b_pass = 0.5, 0.5
expected = a_pass * b_pass
expected += (1 - a_pass) * (1 - b_pass)
print(f"Expected: {expected:.2f}")
print("Kappa:", kappa(0.75, expected))
print("Constant labels:", kappa(1, 1))
With expected agreement 0.50 and observed 0.75, kappa is 0.5. Constant identical labels leave kappa undefined.
Report counts and disagreement examples with a metric. The documentation link explains the formula; this lesson does not require installing scikit-learn.
Predict, explain, then test
Answer these three self-reviewed checks before opening the comparisons. They are not scored. Coding checks use unfamiliar inputs under each task’s stated assumptions.
Count both classes
Do two matching fail labels contribute to observed agreement?
Compare your answer · self-reviewed
Yes. Agreement counts equality for either class, not only pass-pass matches.
Check undefinedness
Expected and observed agreement are both one. What is kappa here?
Compare your answer · self-reviewed
Undefined, represented by None, because 1 - expected is zero. Perfect agreement alone does not make kappa one.
Interpret a negative
Observed is below marginal expected agreement. Must negative kappa be clipped?
Compare your answer · self-reviewed
No. Negative is meaningful under this chance model. Explain the counts, assumptions and disagreements instead.
Choose an exercise to load its prompt.
Loading exercises…
Noted on your route map, with the step you were on. Nothing is lost by parking it — the next review day will bring this idea back, and the flag tells the course where to slow down.
The tutor mounts here when JavaScript is available. The lesson above stays readable without it.