Good place to stop if you're short on time — step 3 picks up here.
Lesson idea Tap to fold
Why might a lead inspect a sample instead of every row?
Show the answer
To audit quality efficiently when a full review is too costly or slow.
Reproduce a sample, not a quality claim
random.sample(population, k) selects occurrences without replacement and leaves the original list unchanged. With unique input IDs, the result has no repeated IDs. k cannot exceed the population size. random.Random(7) creates a private seeded generator; random.seed(7) instead resets the module generator. Repeatability here requires the same inputs, input order, call sequence and runtime version. A seed is neither a security tool nor proof of a fair sample.
The table task uses df.groupby("model", group_keys=False).sample(n=1, random_state=7). It takes one row from each present model group, then sort_values("model") orders the report. Each group needs at least n rows when sampling without replacement. This is stratified coverage: equal numbers per group do not preserve their proportions in the full dataset.
The list tasks use standard Python; the table task loads pandas in the browser. The pinned runtime is Python 3.12 with pandas 2.2.3 and NumPy 2.0.2. Keep a record of the sampling rule and population, and inspect disagreements rather than claiming universal quality from a tiny sample.
# Repeat the same sampling rule on the same IDs.
import random
ids = ["Museum", "Canal", "Plaza", "Rooftop"]
first = random.Random(7).sample(ids, 2)
second = random.Random(7).sample(ids, 2)
print(first == second)
print(len(first), len(set(first)))
print(len(ids))
The independently seeded generators agree, two unique IDs are selected, and the four-item population remains intact.
Choose the sampling rule before looking at outcomes. One row per group helps coverage, but cannot estimate overall prevalence without accounting for group sizes.
Official reference for this lesson · Python sample and seed reference
Predict, explain, then test
Answer these three self-reviewed checks before opening the comparisons. They are not scored. Coding checks use unfamiliar inputs under each task’s stated assumptions.
Explain a seed
What must stay the same to reproduce this seeded sample?
Compare your answer · self-reviewed
The ordered population, sampling rule, call sequence and runtime version. A seed alone does not promise a representative or secure sample.
Check feasibility
Can sample draw three without replacement from two occurrences?
Compare your answer · self-reviewed
No; random.sample raises ValueError when k exceeds the population size.
Compare designs
Does one row per model preserve the full dataset’s model proportions?
Compare your answer · self-reviewed
Not usually. It provides group coverage, not an automatically population-weighted estimate.
Choose an exercise to load its prompt.
Loading exercises…
Noted on your route map, with the step you were on. Nothing is lost by parking it — the next review day will bring this idea back, and the flag tells the course where to slow down.
The tutor mounts here when JavaScript is available. The lesson above stays readable without it.