Good place to stop if you're short on time — step 3 picks up here.
What does json.loads() return for a JSON object?
A Python dictionary.
How do you safely read an optional dictionary field?
Use record.get("field", default).
What should a parser return when a row cannot be trusted?
A clear failure value such as None, or a structured error.
Does checking required-key presence validate the whole record schema?
No. Presence does not check each value’s type or range, or whether required text is empty. Those need separate schema rules.
Record which ideas needed assistance and revisit their lessons. Self-review is not a scored correct answer.
Trace, rebuild and repair controlled records, then independently audit JSONL under the explicit typed policy. Keep rejection reasons and counts; explain the result.
If the rebuild is hard, that is information — not failure.
Lesson idea Tap to fold
Name three reasons to reject a human-data record.
Show the answer
For example: missing id, unknown label, invalid score, malformed JSON, or duplicate id.
Audit a batch without silently losing rows
The original trace converts string/integer scores for a controlled dictionary; the rebuild checks raw nonempty IDs and string labels; the repair handles missing integer scores. Their assumptions differ. They are demonstrations, not one production schema. The independent task combines the ideas under a new, explicit policy.
enumerate(lines, start=1) supplies each line’s number and value. continue moves to the next line. A set stores IDs already accepted; membership checks detect repeats, and add records a newly accepted ID. Create the set inside the function so a second call starts fresh.
For decoded data, check types before operations. Booleans are subclasses of int in Python, so an integer-only score policy must explicitly exclude bool. Catch the expected JSONDecodeError at decoding, then report each first validation failure instead of hiding all errors. The 59-minute practice sequence may be split before the independent task; the quiz is separate.
# Audit a batch without silently losing rows.
import json
text = '{bad}\n{"item_id":"Museum"}\n '
for number, line in enumerate(text.splitlines(), start=1):
if not line.strip():
print(number, "ignored blank")
continue
try:
row = json.loads(line)
except json.JSONDecodeError:
print(number, "json error")
continue
print(number, row.get("item_id"))
1 json error, 2 Museum, then 3 ignored blank. This trace teaches decoding and line numbers; it does not implement the full independent schema.
Use the independent brief for every acceptance and rejection rule. Return an auditable result and check a second batch; the quiz score stays separate.
Predict, explain, then test
These three brief checks are self-reviewed, not a scored quiz. Answer before opening the comparison. Coding checks follow the stated task assumptions on unfamiliar inputs.
Separate types
Why reject the JSON boolean true as an integer score?
Compare your answer · self-reviewed
It decodes to True, which is an int subclass in Python. The explicit schema excludes booleans.
Trace duplicates
Does an invalid row with ID Museum reserve that ID?
Compare your answer · self-reviewed
No. Only a fully valid accepted row reserves an ID; the first later valid row may still pass.
Keep evidence
Why return line numbers and first rejection reasons rather than quietly drop bad rows?
Compare your answer · self-reviewed
They make exclusions inspectable and reproducible. A record count alone cannot explain missing data.
Independent transfer: audit a JSONL batch
Write audit_jsonl(text) in your own implementation. This new policy is stricter than the original demonstration helpers:
- Accept a string argument; raise TypeError for a different caller type.
- Visit text.splitlines() in order, using one-based line numbers. Count whitespace-only lines as ignored. A trailing line ending creates no extra line.
- Decode each nonblank line with json.loads. Catch JSONDecodeError for malformed text and report reason json. A decoded non-dictionary reports record.
- Require item_id to be a string whose stripped value is nonempty. Store that stripped ID; otherwise report item_id.
- Allow only the exact string labels chosen and rejected; otherwise report label.
- Require score to be an integer from 1 through 5, excluding booleans. Do not convert numeric strings or decimals; otherwise report score.
- Keep the first valid occurrence of a trimmed ID. A later valid occurrence reports duplicate. Invalid rows do not reserve IDs.
- Apply those rejection reasons in the order above; report only the first reason for a line. Return {records: [...], errors: [{line: n, reason: ...}], ignored: n}, using Python string keys. Accepted records contain only item_id, label and score; extra input fields are ignored. Preserve order, reset state on each call and return data rather than only printing.
Decoding uses Python json.loads: repeated property names use its last value. Duplicate item IDs follow the separate first-valid-record rule above. JSON syntax and this course schema are different checks.
Select the task in the exercise controls. Its unfinished starter should fail until you implement the rules. The checkpoint’s practical link selects the task and opens this brief without running Python or awarding evidence. Review JSONL decoding or narrow error handling.
Read the unfinished starter without running Python
# Decode rows, then enforce the independent schema.
import json
# Implement the stated policy; this is unfinished.
def audit_jsonl(text):
return {"records": [], "errors": [], "ignored": 0}
# Inspect counts, not a copied report string.
sample = (
'{"item_id":" a1 ","label":"chosen","score":5}\n'
'{broken\n'
'{"item_id":"a1","label":"rejected","score":3}\n'
' \n'
)
result = audit_jsonl(sample)
print("Accepted:", len(result["records"]))
print("Rejected:", len(result["errors"]))
print("Ignored:", result["ignored"])
The sample prints Accepted: 1, Rejected: 2, Ignored: 1 on separate lines. Explain why the second a1 is rejected and why an invalid earlier row would not reserve an ID. The checks inspect returned data for new batches and repeated calls. Quiz results and practical evidence remain separate.
Choose an exercise to load its prompt.
Loading exercises…
Noted on your route map, with the step you were on. Nothing is lost by parking it — the next review day will bring this idea back, and the flag tells the course where to slow down.
Section 5 checkpoint
Answer five questions, review explanations and use the linked refreshers. Your score does not award the independent JSONL audit; tutor login is not required.
The tutor mounts here when JavaScript is available. The lesson above stays readable without it.