DVPPython Studio
Module 8: Model the evaluation domain / Build 4 of 4

Ship a dataset toolkit with evidence

Compose a real import/validate/audit CLI, reconcile loss and verify newly written artifacts without claiming training eligibility.

Runs on your computer · 90–120 minutes · no paid services

Download practice filesFiles, commands & notes

Without JavaScript, use the step links and keep your files on your computer.

One useful idea

A toolkit connects reviewed operations behind one explicit command interface. This build uses the actual Module 6 conversion and provenance helpers and Module 7 reason-text index. The functions are disclosed assistance, not independently rewritten here. Your task owns orchestration and CLI behavior. The CSV fixture has six logical rows: four converted, two quarantined and one converted tie requiring review. CSV records are not necessarily physical lines because quoted cells can contain line breaks.

read_csv bounds input bytes, row count and header shape, then rejects malformed row widths before creating output. ship_dataset calls conversion and publishes into an exclusively NEW directory under an existing regular parent. Existing, linked, reserved or parent-traversal targets are refused. No overwrite, cleanup of user data or rollback is promised. An unexpected write/read-back fault can leave a partial new folder; retain it for inspection and choose another new output name.

The publication contains source.csv, preferences.jsonl, conversion.json, audit.json, notes.json, data-card.json and manifest.json. Preserve raw source bytes, verify actual JSONL typed read-back and its receipt hash, then compare JSON artifacts using typed serialization parity: true is not integer 1. The manifest covers the six other files, not itself. Verify the actual written bytes, not merely a planned hash or an in-memory report. Source and quarantine can contain private text; keep the bundle local until sharing review.

CLI import takes source and --out; validate checks every physical JSONL line; audit additionally requires --manifest covering that dataset in its own root. Status 0 means the operation’s schema checks completed without its declared review conditions, not approved training data. Status 1 means a completed report requires review or provenance failed. Expected input/I/O refusal returns 2 with controlled stderr and no success report; unexpected programming errors remain visible.

The data card states training_eligibility=not_assessed. Ties, quarantined rows, source rights, privacy, reviewer authenticity and fitness require separate human work. Reason-text mentions include negation and do not verify media defects. The browser deliberately has no Build 4 export simulation: execute this on your computer. The temporary worked example owns only its own demo folder; your portfolio task needs persistent evidence and observed defense, not a temporary reference run or quiz score.

Refresh first: Policy versus judgment, Conversion and provenance, Reason-text mentions.

Trace a finished example

import json
from pathlib import Path
from tempfile import TemporaryDirectory
from evalkit.core import ship_dataset
from evalkit.io import audit_with_manifest

source = Path("fixtures/annotations.csv")
before = source.read_bytes()
with TemporaryDirectory(prefix="dvp-owned-example-") as temporary:
    output = Path(temporary) / "new-bundle"
    receipt = ship_dataset(source, output)
    counts = receipt["counts"]
    print(counts["input"], counts["converted"], counts["quarantined"], counts["needs_review"], receipt["exit_status"])
    audit = audit_with_manifest(output / "preferences.jsonl", output / "manifest.json")
    print(len(list(output.iterdir())), audit["provenance"]["checked"])
    card = json.loads((output / "data-card.json").read_text(encoding="utf-8"))
    print(card["training_eligibility"], source.read_bytes() == before)

The real publisher writes seven files into a newly owned temporary demo folder. Six are covered by the manifest; the manifest does not hash itself. Conversion reconciles all six input rows and review status stays 1. The original CSV remains byte-identical. TemporaryDirectory cleans only this example’s own folder on exit; persistent learner bundles are not deleted.

The finished implementation is in evalkit/core.py and evalkit/cli.py. Reading it is guided practice, not independent evidence.

Predict review status

Four rows convert successfully but one is a tie. Is the dataset training-ready?

Compare your answer · self-reviewed

No. A tie remains needs_review, quarantine is accounted for and training_eligibility stays not_assessed. Successful serialization cannot establish rights, privacy, authorship or fitness.

Find false provenance

Is hashing the intended JSON string enough to claim the file was verified?

Compare your answer · self-reviewed

No. Read actual written files, compare typed values and hashes, and verify a manifest that covers this dataset under the same root. A planned hash is not proof of publication.

Recall fault ownership

A write fault leaves a partial NEW directory. Should the CLI silently delete it or announce success?

Compare your answer · self-reviewed

Neither. Preserve the partial directory, surface the fault and choose another new output name. This workflow promises no overwrite or rollback; unexpected programming failures stay visible.

Change it, then build your own

One controlled change

Import the fixture into your own NEW folder and run validate and audit with its manifest. Modify a copied output file, rerun audit and explain the changed provenance status. Never edit the supplied fixtures or reuse an existing destination.

Your independent task

Implement ship_dataset in practice.py using disclosed read_csv, conversion and publish_bundle helpers, then implement main(argv=None) in practice_cli.py with import/validate/audit subcommands. Use your own practice importer, require --out/--manifest where declared, print controlled actual reports and return their statuses. Catch only expected refusal classes; unexpected faults must propagate. Keep the entry guard. Do not import the finished CLI.

What success looks like

Build 4 checks real writes, source-byte retention, typed round-trips, six-file provenance, loss/review reconciliation, CLI statuses and no overwrite. The full reference suite has 88 checks across all builds; it is not proof you completed your own stubs. Keep a persistent changed-input bundle and the Portfolio II evidence template for separate review.

Hint 1 · a question

List the three commands, required paths and expected 0/1/2 meanings. Which helper verifies actual files rather than merely describing them?

Hint 2 · a concept cue

Compute the operation’s controlled report first, print it only after completion, then return its exit_status. Keep input refusal distinct from unexpected programming faults.

Hint 3 · a localized example

if __name__ == "__main__": raise SystemExit(main()) makes script execution call your CLI without launching it on import. Your importer must use practice.ship_dataset, not the finished reference CLI.

Need the complete worked solution?

Open evalkit/core.py and evalkit/cli.py from the kit. Trace it, close it, then try fresh inputs in your own files. Treat the attempt as guided; seeing the solution does not award a practical pass.

Course help is guidance, not independent evidence. With JavaScript, opening help records guidance locally; otherwise note it in your README. Reset does not erase that history.

Repair a failed check

If the command exits silently, check the __main__ guard. If valid imports always return 0, return the actual review status. If malformed input produces a folder or success output, validate before publication/output. If bool and integer JSON compare equal, use typed parity. If a source or write fault is hidden, remove the broad fallback.

NotImplementedError means a practice stub is still unfinished. Read the failing test name and the last error line. Change one behavior, rerun that build, then rerun all implemented builds.

Show it works on new inputs

Create a changed six-row invented CSV, run your own importer into a persistent NEW folder, reproduce conversion/review counts and audit all six manifest entries. Keep commands, statuses, tests, data-card limitations and helper disclosure in PORTFOLIO.md. Explain a corrupted output and a refused destination during a separate observed defense; no automated check awards that defense.

Self-review: name the input, result, refused case and reason. Your local test output and explanation are separate from a quiz score; this page does not certify a pass.

Keep the idea

Ship artifacts that another person can inspect, reproduce and question. A verified file bundle remains distinct from eligible training data and independent portfolio work.

Organize your portfolio evidence and fresh defense. Keep local files, assistance and observed gaps separate from this quiz.

Module 8 checkpoint

Five questions, followed by the separate practical task above. JavaScript loads the scored questions; the build, files and hints remain available without it.