Good place to stop if you're short on time — step 3 picks up here.
What does a comparison such as score >= 4 produce?
A Boolean value—True or False—for each value being compared.
Why can a missing label be more serious than a missing note?
The label may be required for the analysis; a note may be optional.
What two numbers define a rate?
A count of matching cases and the total eligible cases.
For grouped scores with missing values, what do size, count and mean measure?
size counts all rows in the group; score count counts present scores; score mean averages only present scores. A group with no scores has a missing mean, not zero.
Record which ideas needed assistance and revisit their lessons. Self-review is not a scored correct answer.
Trace, rebuild and repair table operations, then independently audit rows under an explicit cleaning policy. Report exclusions and denominators; explain what the retained sample can and cannot show.
If the rebuild is hard, that is information — not failure.
Lesson idea Tap to fold
What can a duplicate item do to an acceptance rate?
Show the answer
It gives one decision extra weight and can distort the result.
Make cleaning order and exclusions inspectable
Repeated item IDs can overweight a summary. duplicated("item_id") marks later occurrences; drop_duplicates("item_id", keep="first") retains the first row in the current order. This is a policy, not proof that the first observation is the most accurate.
Order matters. Drop rows missing required facts and reject invalid labels before deduplicating, so an unusable earlier row does not reserve an ID. Work on a new table; do not alter the caller’s original evidence. copy(deep=True) is useful when you need to make separate edits.
A Series.isin(["pass", "fail"]) call builds a boolean mask for membership in the allowed labels. Check required-column membership with a set’s issubset before selecting columns. In a group loop, round(float(value), 2) turns an observed numeric mean into a two-decimal Python value; only call it when the scored subset is nonempty.
The independent function raises TypeError for a wrong caller type and ValueError for missing required columns. It assumes upstream typed values, reports sequential exclusions and returns per-model denominators. groupby supplies (model, group) pairs to a loop; len(group) counts rows and dropna on score gives the scored subset. Return None for an unscored mean. Allow about 60 minutes for the four tasks, or split before the independent audit; the quiz is separate.
# Make cleaning order and exclusions inspectable.
import pandas as pd
df = pd.DataFrame([
{"item_id": "Museum", "label": None},
{"item_id": "Museum", "label": "pass"},
{"item_id": "Museum", "label": "fail"},
])
usable = df.dropna(subset=["label"])
clean = usable.drop_duplicates("item_id", keep="first")
print("Raw:", len(df), "Usable:", len(usable))
print(clean["label"].tolist())
Raw: 3 Usable: 2, then ['pass']. The missing first label is removed before the first valid observation is retained.
Use the independent brief to define counts and denominators, then defend the result on unfamiliar data. A quiz score is separate from this practical evidence.
Predict, explain, then test
These three brief checks are self-reviewed, not scored. Answer before opening the comparison. Coding checks use unfamiliar tables under the stated task assumptions.
Trace cleaning order
A missing-label ID appears before a valid row with the same ID. Which should reserve the ID?
Compare your answer · self-reviewed
The first remaining valid row, after the required-field/label exclusions.
Preserve an empty mean
A retained group has two rows but no observed scores. What are scored and mean_score?
Compare your answer · self-reviewed
scored is zero and mean_score is None; the two labelled rows still support its acceptance rate.
Keep evidence
Why return exclusions and avoid changing the caller’s table?
Compare your answer · self-reviewed
Counts make the retained sample explainable; preserving the original permits checking or revising the policy later.
Independent transfer: a defensible dataset summary
Write summarize_dataset(frame) independently. Use this explicit policy rather than treating cleaning as an invisible step:
- Accept a pandas DataFrame; raise TypeError for another caller type. Require item_id, model, label and score columns; raise ValueError if any are absent, including for an empty table.
- This task assumes typed upstream data: IDs/models are nonempty strings or missing, labels are strings or missing, and scores are finite numeric 1–5 or missing. It is a table audit, not a replacement for Day 25’s untrusted-record schema. Extra columns are allowed.
- Count raw_rows. Drop rows missing item_id, model or label; count them as missing_required. A missing score alone does not reject a row.
- Of the remaining rows, keep only labels pass and fail; count other labels as invalid_label.
- Then remove duplicate item_id values globally, keeping the first remaining row in input order. Count removed rows as duplicates and remaining rows as kept. Earlier missing/invalid rows do not reserve IDs.
- For each retained model, return rows, scored, mean_score and acceptance_rate. scored counts only present scores; mean_score is their mean rounded to two decimals, or None when none are present. acceptance_rate is pass rows divided by all retained rows for that model, rounded to two decimals.
- Return a dictionary with raw_rows, missing_required, invalid_label, duplicates, kept and models (a model-name dictionary of those four measures). Counts are integers. Do not mutate the input, fill missing scores with zero or hide caller errors. Empty valid input returns zero counts and an empty models dictionary.
Select the task in the exercise controls. Its unfinished starter should fail until you implement the rules. The quiz’s practical link selects it and opens this brief without running code or awarding evidence. Review missingness and group denominators.
Read the unfinished starter without running Python
# Audit the table under the independent policy.
import pandas as pd
# Implement the rules; this placeholder must fail.
def summarize_dataset(frame):
return {"kept": 0, "duplicates": 0, "models": {}}
# Trace duplicate order and missing-score policy.
sample = pd.DataFrame([
{"item_id": "a1", "model": "Museum",
"label": "pass", "score": 5},
{"item_id": "a1", "model": "Museum",
"label": "fail", "score": 1},
{"item_id": "a2", "model": "Canal",
"label": "fail", "score": None},
{"item_id": "a3", "model": "Canal",
"label": None, "score": 3},
])
result = summarize_dataset(sample)
print("Kept:", result["kept"])
print("Duplicates:", result["duplicates"])
The sample prints Kept: 2 and Duplicates: 1 on separate lines. Museum has one scored retained row; Canal has one retained row but no observed score. Explain why Canal’s mean is None while its labelled row still contributes to its acceptance rate. Checks inspect returned data on new, empty and nullable tables, repeated calls and unchanged inputs. Quiz scores and practical evidence stay separate.
Choose an exercise to load its prompt.
Loading exercises…
Noted on your route map, with the step you were on. Nothing is lost by parking it — the next review day will bring this idea back, and the flag tells the course where to slow down.
Section 6 checkpoint
Answer five questions, review explanations and use the linked refreshers. Your score does not award the independent dataset audit; tutor login is not required.
The tutor mounts here when JavaScript is available. The lesson above stays readable without it.