DVPPython Studio
Module 3: Process the whole bin / Build 2 of 4

Summarize generation batches

Count all records, missing metadata and duplicate identities without quietly changing the denominator.

Runs on your computer · 45–65 minutes · no paid services

Download practice filesFiles, commands & notes

Without JavaScript, use the step links and keep your files on your computer.

One useful idea

Grouping turns repeated labels into counts. collections.Counter can count a sequence of labels; a dictionary with get(label, 0) + 1 is an equivalent approach. It does not decide which records you should remove. Here every supplied row counts, including repeated IDs.

Require exactly id/provider/model/status. Strip text labels; present labels need 1–80 characters without remaining control characters. Provider/model may be None or blank, counted separately as missing. IDs/status cannot be missing. Status is casefolded to pass/review/block; provider/model grouping is case-sensitive and does not guess aliases. These are invented fixture categories, not current vendor measurements.

Return total rows, distinct cleaned IDs, extra duplicate occurrences, sorted repeated IDs, present-label counts, all three status counts and separate provider/model missing counts. Each known-label total plus its own missing count equals the row total. A duplicate warning does not authorize silently dropping a record, and a malformed caller row is an error rather than an exclusion.

Refresh first: Cleaned lookup keys and preserved evidence, Dictionaries, Counts and denominators.

Trace a finished example

python -X utf8 demo.py --build 2
# Rows: 25; Unique IDs: 24; Duplicate rows: 1; Duplicate IDs: gen-03
# by_status: pass 17, review 5, block 3; missing provider 2, model 2

The reference validates and counts 25 rows. gen-03 appears twice, so unique IDs are 24 and duplicate rows are 1; both status observations remain. The demonstration formats the returned counts into an aligned table only after the function succeeds.

The finished implementation is in collection_tools.py. Reading it is guided practice, not independent evidence.

Predict the denominator

25 rows contain one repeated ID. How many rows count in the status totals?

Compare your answer · self-reviewed

25. There are 24 unique IDs, but this summary deliberately counts every recorded observation.

Find the invented label

Why not group missing models under a known provider name?

Compare your answer · self-reviewed

That invents metadata. Count missing model separately; known model counts plus that missing count reconcile to all rows.

Explain an empty result

What should an empty batch report?

Compare your answer · self-reviewed

Zero rows/unique/duplicate totals, empty provider/model groups, zero values for all three statuses and zero missing counts. No rate or made-up group is needed.

Try the idea in this browser

Runs in this browser · optional preparation · local project checks remain separate

Try a small function before opening your local files. Python downloads when you choose Run; if it cannot load, your code stays here and the local kit still works. The worker executes on your device, not on a DVP server. Only run code you trust: this is not a hostile-code security sandbox.

JavaScript loads the practice controls. Python starts only after Run.

Read the browser task briefs without running Python

Trace collection identity

Read the complete function, predict a changed case, then Run. Implement summarize_batches(rows) with the exact README fields and label policy. Count ALL rows, distinct cleaned IDs, extra duplicates and sorted repeated IDs. Count present provider/model labels and missing optional metadata separately; include pass/review/block zeros. Preserve input and use fresh output. Wrong types raise TypeError; malformed keys/labels/status raise ValueError. These are fixture observations, not provider measurements.

Preserve the changed collection independently

Implement summarize_batches(rows) with the exact README fields and label policy. Count ALL rows, distinct cleaned IDs, extra duplicates and sorted repeated IDs. Count present provider/model labels and missing optional metadata separately; include pass/review/block zeros. Preserve input and use fresh output. Wrong types raise TypeError; malformed keys/labels/status raise ValueError. These are fixture observations, not provider measurements.

Change it, then build your own

One controlled change

In a copied batch fixture, add a second repeated ID and a missing model. Predict row/unique/duplicate/missing counts before running the summary; do not deduplicate the fixture.

Your independent task

Implement summarize_batches(rows) in practice.py under the exact README schema and label policy. Return exactly rows, unique_ids, duplicate_rows, duplicate_ids, by_provider, by_model, by_status and missing. Count all rows, sort repeated IDs, keep status zeros, preserve input and create fresh outputs. Wrong types are TypeError; keys/values/status errors are ValueError.

What success looks like

Four summary test methods pass, including 25 and 27 changed records, missing metadata, repeated cleaned IDs, native integer counts, empty input, caller errors and preserved input. The supplied 25-row fixture also produces the demonstrated aligned report.

Hint 1 · a question

Write separate definitions for row count, unique identity count and extra duplicate occurrences.

Hint 2 · a concept cue

Count cleaned IDs and present labels separately; increment missing counts without inventing keys. Keep pass/review/block zeros even when no row has that status.

Hint 3 · a localized example

duplicate_rows is len(rows) minus the number of distinct cleaned IDs. duplicate_ids contains sorted IDs whose count exceeds 1; it is not the list of retained rows.

Need the complete worked solution?

Open collection_tools.py from the kit. Trace it, close it, then try fresh inputs in your own files. Treat the attempt as guided; seeing the solution does not award a practical pass.

Course help is guidance, not independent evidence. With JavaScript, opening help records guidance locally; otherwise note it in your README. Reset does not erase that history.

Repair a failed check

If status totals equal unique IDs, look for a dictionary that overwrites duplicate rows. If a missing label becomes a known group, remove the fallback. If changing the result changes a later call, create count dictionaries inside the function.

NotImplementedError means a practice stub is still unfinished. Read the failing test name and the last error line. Change one behavior, rerun that build, then rerun all implemented builds.

Show it works on new inputs

Use at least 25 new fixture observations with two repeated IDs and different missing-provider/model positions. Reconcile each grouping total plus missing count to all rows, and explain why a deduplication policy would be a different report.

Self-review: name the input, result, refused case and reason. Your local test output and explanation are separate from a quiz score; this page does not certify a pass.

Keep the idea

A report is defensible when its units, denominators, missingness and duplicate policy are visible.