DynamicVibePython learning
How to study · reference & help

One lesson at a time

Read the idea, then follow the named exercise steps. On a phone, use Idea, Code and Output to switch views. Review days begin with recall before the rebuild. A self-review mark is not a scored answer.

Run, inspect, repair

Run loads Python in your browser on the first use, so you need an internet connection. Read Output and the check result. Stop interrupts a run; it keeps your code. Reset restores the selected starter, replacing that exercise’s draft. Use an authored hint before asking for help.

Keep your work

Code drafts, progress and Notes save on this device when browser storage is available. They do not sync through your login. Use Backup & records above to download saved learning records, preview an import or undo the latest import. Copy unsaved text first; clearing browser data can remove local work. A save warning means you should copy your work before leaving.

Practice files and reports

Each browser Run starts in a fresh temporary workspace. A sample file exists only when that run’s code creates it; these are not your computer’s files. On report tasks, use Preview to inspect the actual text and Download to request a copy before another Run, Reset, exercise change or leaving. Reports do not save through tutor login. Read files and write reports.

Get tutoring when you need it

Lessons and checkpoints do not need tutor login. Unlock the tutor through Training’s shared tutor access, then use its return link to come back here. Your question is only sent when you press Send. Do not paste passwords or private production data.

Python words in plain language

Assignment and type
frames = 96 binds a name to an integer. "96" is text, not the same numeric value. Values and names.
List and index
A list keeps items in order. Positions start at zero; the final nonnegative index is length minus one. List lookups.
Loop and indentation
A for loop processes each item. Indentation groups its repeated statements; a statement outside the block runs separately. Loops.
Parameter, argument and Return
A parameter names an input in a definition; an argument supplies its value in a call. Return gives a result to the caller; print displays it. A function that only prints returns None. Functions.
Dictionary and schema
A dictionary labels values with keys. A schema states the required fields and types; a missing fact is not automatically zero. Dictionaries · Validation.
File, CSV and JSONL
A file stores bytes/text. CSV needs a quoting-aware parser; JSONL contains one JSON value per line. Browser practice files are temporary, not computer files. CSV · JSONL.
Denominator and evidence
A rate divides a count by its declared eligible total. Keep exclusions visible. A quiz score, practical check and self-review are different evidence. Quality rates.

Read the error, then make one repair

SyntaxError / IndentationError
Python could not parse the code. Inspect the named line and the preceding line for a missing colon, unmatched quote/bracket or inconsistent indentation. Review blocks.
NameError
A name has no value here. Compare case and spelling, and check that assignment happens before use. Do not replace an unknown parameter with a fixed sample.
TypeError / ValueError
A wrong kind of value or an unacceptable value can fail an operation. Inspect the supplied input and declared policy; do not broadly suppress the error. Narrow error handling.
IndexError / KeyError
The requested position or key is missing. Inspect the collection and interface instead of guessing a fallback. Indices · Keys.
Run stays loading / runtime or checker unavailable
This is not proof that your answer is wrong. Stop if available, copy your draft, check your connection and retry. If it repeats, keep the visible error and use the readable lesson/hints. Never reset or clear browser storage as the first repair.
Expected sample prints, but a check fails
Read the named contract and test changed/empty inputs. The helper may return None, use a fixed value or stop too early even when the sample looks right. A failed check is not a request to copy the expected output.

Press Escape while working inside this panel to close it and return focus to its summary. Without JavaScript, use the summary again. For local Studio commands, use terminal troubleshooting.

← Route Week 7 · Day 35 · Checkpoint

Evaluator calibration review

Loading today's steps…
Lesson idea Tap to fold
Recall

What should happen after an agreement rate falls below target?

Show the answer

Inspect disagreement examples, clarify the rubric, calibrate, and measure again.

Today

Keep evidence before judging a metric

Rebuild the reference-agreement table using a boolean correct column, then its mean within each evaluator. A strict rate < 0.75 is only a classroom review flag: exactly 0.75 does not trigger it. A rounded display and a small sample do not justify automatically judging or punishing an evaluator. Inspect disagreements and clarify the rubric first.

The independent task combines validation, first-valid retention and confusion counts. Do not globally deduplicate item_id: two different evaluators rating the same item are two observations. A tuple (evaluator, item_id) is an immutable composite key. A set starts with set(); key in seen tests membership and seen.add(key) records a retained pair. Validate and strip string IDs before reserving a key.

Create fresh groups inside each function call, accumulate TP/FP/FN/TN, then compute each denominator. Return data, not just its printed representation. Keep the input unchanged; report raw, rejected, duplicate and retained counts plus disagreement IDs. Return None for undefined rates or kappa, and round defined metrics to four decimals after calculation. Allow about 59 minutes for four tasks; split before the independent report if needed. The quiz is separate.

# Retain pairs, not globally unique item IDs.
pairs = [("Ava", "A"), ("Ava", "A"), ("Ben", "A")]
seen = set()
kept = []
for key in pairs:
    if key in seen:
        continue
    seen.add(key)
    kept.append(key)
print(kept)
print("Same item, separate evaluators:", len(kept))

The duplicate Ava/A pair is removed, while Ben/A is retained. Identity depends on the stated composite key, not just item ID.

Explain Ava’s false-positive row and Ben’s undefined kappa in the independent sample. A passing software check is practical evidence, not a claim of general evaluator quality.

Official reference for this lesson

Predict, explain, then test

Answer these three self-reviewed checks before opening the comparisons. They are not scored. Coding checks use unfamiliar inputs under each task’s stated assumptions.

Choose identity

Ava and Ben rate the same item ID. Should one row be dropped as a duplicate?

Compare your answer · self-reviewed

No. Different evaluator/item pairs are separate observations; only later valid occurrences of the same pair are duplicates.

Inspect the boundary

Does exactly 0.75 trigger the classroom below-75% flag?

Compare your answer · self-reviewed

No. The comparison is strictly less than; use the full rate, not the rounded display.

Preserve evidence

Why return None and disagreement IDs alongside rates?

Compare your answer · self-reviewed

None distinguishes absence from zero; disagreement IDs let a person inspect retained evidence and the rubric before interpreting a metric.

Independent transfer: an evaluator evidence report

Write summarize_evaluators(rows) independently from this explicit report policy:

  1. Write summarize_evaluators(rows). Accept a list, otherwise raise TypeError. Do not change its rows or dictionaries; start fresh on every call.
  2. Reject a row unless it is a dictionary with nonempty string item_id and evaluator after stripping outer whitespace, and exact string gold/label values pass or fail. Count rejected rows separately; do not convert numbers or booleans into labels. Extra fields are allowed.
  3. After validation, keep the first valid (evaluator, item_id) pair. A duplicate for that evaluator counts as duplicates; the same item rated by a different evaluator remains separate. Invalid earlier rows do not reserve a pair.
  4. Return raw_rows, rejected, duplicates, kept and evaluators. evaluators maps each cleaned evaluator name to rows, tp, fp, fn, tn, agreement, precision, recall, kappa and disagreements. Counts are integers; disagreements is an input-order list of retained item IDs where label differs from gold.
  5. Treat pass as the positive class. TP is gold pass / label pass; FP is gold fail / label pass; FN is gold pass / label fail; TN is gold fail / label fail. Agreement is (tp + tn) / rows, precision is tp / (tp + fp), and recall is tp / (tp + fn).
  6. For kappa, observed is agreement. Gold-pass share is (tp + fn) / rows; predicted-pass share is (tp + fp) / rows. expected is their product plus the product of their complements. kappa is (observed - expected) / (1 - expected).
  7. Return None for precision or recall with a zero denominator, and for kappa when expected is one. Round defined metrics to four decimals only after calculation. Empty input returns zero counts and an empty evaluators dictionary. Explain the retained sample and disagreement examples; these rates do not prove general evaluator quality or justify automatic punishment.

Select the independent task in the exercise controls. The quiz’s practical link selects it and opens this brief without running code or awarding evidence. Review validation, precision and recall, and chance-corrected agreement.

Read the unfinished starter without running Python
# Summarize retained evidence, not a leaderboard.
def summarize_evaluators(rows):
    # Implement the policy; this placeholder must fail.
    return {"kept": 0, "evaluators": {}}

# Keep repeated and cross-evaluator IDs distinct.
rows = [
    {"item_id": "a1", "evaluator": "Ava",
     "gold": "pass", "label": "pass"},
    {"item_id": "a2", "evaluator": "Ava",
     "gold": "fail", "label": "pass"},
    {"item_id": "a1", "evaluator": "Ava",
     "gold": "pass", "label": "fail"},
    {"item_id": "a1", "evaluator": "Ben",
     "gold": "pass", "label": "pass"},
    {"item_id": "", "evaluator": "Ava",
     "gold": "fail", "label": "fail"},
]
result = summarize_evaluators(rows)
print("Kept:", result["kept"])
print("Evaluators:", len(result["evaluators"]))

The completed sample prints Kept: 3 and Evaluators: 2 on separate lines. Ava retains two rows: TP 1, FP 1, FN 0, TN 0, agreement 0.5, precision 0.5, recall 1, kappa 0 and disagreements [a2]. Ben retains one TP row: agreement, precision and recall 1; kappa None; no disagreements. Explain why Ben’s perfect observed match does not define kappa. Checks inspect returned data on new, empty, malformed and duplicate inputs, repeated calls and unchanged rows. Your quiz score and practical evidence stay separate.

Now

Choose an exercise to load its prompt.

day-35.py Python · packages shown per task · runs here
Output

Loading exercises…


    
    

Section 7 checkpoint

Answer five questions, review explanations and use the linked refreshers. The independent evaluator report is a separate practical task; neither requires tutor login.

AI Tutor

The tutor mounts here when JavaScript is available. The lesson above stays readable without it.