DVPPython Studio
Module 6: Shape trustworthy data / Build 4 of 4

Convert annotations safely

Map declared labels, preserve source lineage, quarantine refusals and account for named conversion losses.

Runs on your computer · 70–100 minutes · no paid services

Download practice filesFiles, commands & notes

Without JavaScript, use the step links and keep your files on your computer.

One useful idea

Conversion is a contract between formats, not permission to discard unfamiliar data. Source fields are annotation_id, prompt, asset_a, asset_b, label, reason, evaluator and created_at. Extra JSON fields are accepted only as named conversion losses. A/B/tie is the declared vocabulary after strip/casefold lookup; preserve the ORIGINAL source_label in lineage.

A chooses asset_a; B swaps the pair; tie preserves source ordering with needs_review, never a fabricated training winner. Every converted preference must pass the exact schema. The first accepted ID owns identity; later valid duplicates are quarantined. Unknown labels use unknown_label, invalid/missing fields schema, repeated accepted IDs duplicate. A prior refused row does not consume a usable ID.

The report has converter_version, input_rows, converted_rows, quarantined_rows, dropped_fields, records, quarantine and lineage. Each quarantine has row, controlled code and a fresh complete source copy. Lineage has record_id, row and original source_label. Reconcile input=converted+quarantined. Count dropped fields only on converted rows: quarantined sources retain those values. Counts name losses but cannot reconstruct dropped VALUES; retain the source CSV plus byte manifest. Quarantine can contain private supplied text; never send it to tutor logs automatically.

Sources must be losslessly round-trip JSON data, finite, at most 65,536 encoded bytes and within strict depth/number limits. Unretainable objects, key/tuple rewrites, invalid Unicode or oversized data refuse the WHOLE operation, not fake a reconciled source. At most 100,000 rows; this report holds records/quarantine/lineage, so it is not constant-memory streaming. The CSV demo uses UTF-8 and newline="", preserves quoted commas, treats values as strings and refuses missing/duplicate headers or mismatched widths. It is one authored migration, not a universal annotation adapter or completed Dataset Toolkit.

Refresh first: Preference schema and review-needed verdict, Input accounting and explicit refusals, Source-byte provenance.

Trace a finished example

from pathlib import Path
from demo import read_annotations
from dataset_tools.core import convert_annotations

result = convert_annotations(read_annotations(Path("fixtures/annotations.csv")))
print(result["input_rows"], result["converted_rows"], result["quarantined_rows"])
print(result["dropped_fields"])
print([item["code"] for item in result["quarantine"]])
print([item["source_label"] for item in result["lineage"]])

Six real CSV fixture rows enter. A/B/tie become three validated records; unknown, duplicate and same-asset rows remain in source-preserving quarantine. Confidence is dropped on only the three converted rows. Original label casing survives in lineage, and input six equals three plus three.

The finished implementation is in dataset_tools/core.py and demo.py. Reading it is guided practice, not independent evidence.

Predict losses

Confidence exists in all six fixture rows. Why is its dropped count three?

Compare your answer · self-reviewed

Only three converted rows lose that field. The other three retain confidence in their full quarantine source, so counting six would exaggerate actual loss.

Find invented meaning

Should an unknown label default to A to keep counts simple?

Compare your answer · self-reviewed

No. Quarantine unknown_label and retain source meaning. An invented winner is not a safe migration or useful training evidence.

Recall source evidence

Can the loss counts reconstruct confidence values?

Compare your answer · self-reviewed

No. Counts describe what was lost, not the values. Keep source data and a provenance manifest; retain lineage separately from the converted preference.

Try the idea in this browser

Runs in this browser · optional preparation · local project checks remain separate

Try a small function before opening your local files. Python downloads when you choose Run; if it cannot load, your code stays here and the local kit still works. The worker executes on your device, not on a DVP server. Only run code you trust: this is not a hostile-code security sandbox.

JavaScript loads the practice controls. Python starts only after Run.

Read the browser task briefs without running Python

Guided annotation conversion

Reviewed preference validation, bounded JSON copying and parser helpers are supplied and disclosed. Write your own conversion/accounting workflow; the finished converter is not supplied. Use only invented in-memory rows here. No file, CSV read, JSONL stream, byte receipt or provenance operation runs. Preserve source data and reconcile accepted/quarantined rows; needs_review is unresolved, not a training winner. This is browser preparation, not local project or Portfolio II evidence. Implement convert_annotations in practice.py. Reviewed preference/model and fresh bounded-source helpers may be reused and disclosed, but implement your own declared mapping, accepted-ID policy, quarantine, original-label lineage and named losses. Return fresh data without mutating source; refuse unretainable inputs rather than invent a source report. Rerun every implemented practice group and keep the original CSV separate.

Independent annotation conversion

Reviewed preference validation, bounded JSON copying and parser helpers are supplied and disclosed. Write your own conversion/accounting workflow; the finished converter is not supplied. Use only invented in-memory rows here. No file, CSV read, JSONL stream, byte receipt or provenance operation runs. Preserve source data and reconcile accepted/quarantined rows; needs_review is unresolved, not a training winner. This is browser preparation, not local project or Portfolio II evidence. Implement convert_annotations in practice.py. Reviewed preference/model and fresh bounded-source helpers may be reused and disclosed, but implement your own declared mapping, accepted-ID policy, quarantine, original-label lineage and named losses. Return fresh data without mutating source; refuse unretainable inputs rather than invent a source report. Rerun every implemented practice group and keep the original CSV separate.

Change it, then build your own

One controlled change

Add new invented B and tie rows with a spaced/capitalized label, plus one new unknown label and extra field. Predict pair order, review verdict, original lineage and exact loss/reconciliation counts.

Your independent task

Implement convert_annotations in practice.py. Reviewed preference/model and fresh bounded-source helpers may be reused and disclosed, but implement your own declared mapping, accepted-ID policy, quarantine, original-label lineage and named losses. Return fresh data without mutating source; refuse unretainable inputs rather than invent a source report. Rerun every implemented practice group and keep the original CSV separate.

What success looks like

The build4 group checks A/B/tie semantics, source preservation, exact row reconciliation/loss counts, unknown/malformed/duplicate behavior, fresh ownership and whole-operation refusal for unretainable JSON data. The actual CSV demo agrees with those counts. This does not complete Portfolio II or establish automatic training eligibility.

Hint 1 · a question

Write the output and quarantine for A, B, tie, unknown and duplicate before coding. Which source facts must survive?

Hint 2 · a concept cue

Validate a candidate preference, reserve only accepted IDs, retain fresh refused sources and count extra keys only when conversion drops them.

Hint 3 · a localized example

For B, chosen_id comes from asset_b and rejected_id from asset_a. A lookup may use label.strip().casefold(), but source_label must retain the original string.

Need the complete worked solution?

Open dataset_tools/core.py and demo.py from the kit. Trace it, close it, then try fresh inputs in your own files. Treat the attempt as guided; seeing the solution does not award a practical pass.

Course help is guidance, not independent evidence. With JavaScript, opening help records guidance locally; otherwise note it in your README. Reset does not erase that history.

Repair a failed check

If B still chooses asset_a, trace the explicit swap. If unknown labels become winners, remove the default. If duplicate rows disappear, retain quarantine and physical row identity. If loss counts include quarantined extras, count only converted losses. If source labels lose case/spaces, separate lookup text from original lineage.

NotImplementedError means a practice stub is still unfinished. Read the failing test name and the last error line. Change one behavior, rerun that build, then rerun all implemented builds.

Show it works on new inputs

Supply a changed six-row source with A/B/tie/unknown/duplicate/invalid cases and one extra field. Show before/after samples, original lineage, named losses and input=converted+quarantined. Keep source bytes/manifest, explain review eligibility and disclose helper/reference assistance.

Self-review: name the input, result, refused case and reason. Your local test output and explanation are separate from a quiz score; this page does not certify a pass.

Keep the idea

A trustworthy conversion exposes lost and unresolved meaning. Useful counts never justify invented labels or silent data removal.

Module 6 checkpoint

Five questions, followed by the separate practical task above. JavaScript loads the scored questions; the build, files and hints remain available without it.