One useful idea
Conversion is a contract between formats, not permission to discard unfamiliar data. Source fields are annotation_id, prompt, asset_a, asset_b, label, reason, evaluator and created_at. Extra JSON fields are accepted only as named conversion losses. A/B/tie is the declared vocabulary after strip/casefold lookup; preserve the ORIGINAL source_label in lineage.
A chooses asset_a; B swaps the pair; tie preserves source ordering with needs_review, never a fabricated training winner. Every converted preference must pass the exact schema. The first accepted ID owns identity; later valid duplicates are quarantined. Unknown labels use unknown_label, invalid/missing fields schema, repeated accepted IDs duplicate. A prior refused row does not consume a usable ID.
The report has converter_version, input_rows, converted_rows, quarantined_rows, dropped_fields, records, quarantine and lineage. Each quarantine has row, controlled code and a fresh complete source copy. Lineage has record_id, row and original source_label. Reconcile input=converted+quarantined. Count dropped fields only on converted rows: quarantined sources retain those values. Counts name losses but cannot reconstruct dropped VALUES; retain the source CSV plus byte manifest. Quarantine can contain private supplied text; never send it to tutor logs automatically.
Sources must be losslessly round-trip JSON data, finite, at most 65,536 encoded bytes and within strict depth/number limits. Unretainable objects, key/tuple rewrites, invalid Unicode or oversized data refuse the WHOLE operation, not fake a reconciled source. At most 100,000 rows; this report holds records/quarantine/lineage, so it is not constant-memory streaming. The CSV demo uses UTF-8 and newline="", preserves quoted commas, treats values as strings and refuses missing/duplicate headers or mismatched widths. It is one authored migration, not a universal annotation adapter or completed Dataset Toolkit.
Refresh first: Preference schema and review-needed verdict, Input accounting and explicit refusals, Source-byte provenance.
Trace a finished example
from pathlib import Path
from demo import read_annotations
from dataset_tools.core import convert_annotations
result = convert_annotations(read_annotations(Path("fixtures/annotations.csv")))
print(result["input_rows"], result["converted_rows"], result["quarantined_rows"])
print(result["dropped_fields"])
print([item["code"] for item in result["quarantine"]])
print([item["source_label"] for item in result["lineage"]])Six real CSV fixture rows enter. A/B/tie become three validated records; unknown, duplicate and same-asset rows remain in source-preserving quarantine. Confidence is dropped on only the three converted rows. Original label casing survives in lineage, and input six equals three plus three.
The finished implementation is in dataset_tools/core.py and demo.py. Reading it is guided practice, not independent evidence.
Predict losses
Confidence exists in all six fixture rows. Why is its dropped count three?
Compare your answer · self-reviewed
Only three converted rows lose that field. The other three retain confidence in their full quarantine source, so counting six would exaggerate actual loss.
Find invented meaning
Should an unknown label default to A to keep counts simple?
Compare your answer · self-reviewed
No. Quarantine unknown_label and retain source meaning. An invented winner is not a safe migration or useful training evidence.
Recall source evidence
Can the loss counts reconstruct confidence values?
Compare your answer · self-reviewed
No. Counts describe what was lost, not the values. Keep source data and a provenance manifest; retain lineage separately from the converted preference.
Try the idea in this browser
Runs in this browser · optional preparation · local project checks remain separate
Try a small function before opening your local files. Python downloads when you choose Run; if it cannot load, your code stays here and the local kit still works. The worker executes on your device, not on a DVP server. Only run code you trust: this is not a hostile-code security sandbox.
JavaScript loads the practice controls. Python starts only after Run.
The tutor button only prepares a question locally. Review it and choose Send yourself; no code is sent merely by running or opening a lesson.
Output
Errors and check feedback
Read the browser task briefs without running Python
Guided annotation conversion
Reviewed preference validation, bounded JSON copying and parser helpers are supplied and disclosed. Write your own conversion/accounting workflow; the finished converter is not supplied. Use only invented in-memory rows here. No file, CSV read, JSONL stream, byte receipt or provenance operation runs. Preserve source data and reconcile accepted/quarantined rows; needs_review is unresolved, not a training winner. This is browser preparation, not local project or Portfolio II evidence. Implement convert_annotations in practice.py. Reviewed preference/model and fresh bounded-source helpers may be reused and disclosed, but implement your own declared mapping, accepted-ID policy, quarantine, original-label lineage and named losses. Return fresh data without mutating source; refuse unretainable inputs rather than invent a source report. Rerun every implemented practice group and keep the original CSV separate.
Independent annotation conversion
Reviewed preference validation, bounded JSON copying and parser helpers are supplied and disclosed. Write your own conversion/accounting workflow; the finished converter is not supplied. Use only invented in-memory rows here. No file, CSV read, JSONL stream, byte receipt or provenance operation runs. Preserve source data and reconcile accepted/quarantined rows; needs_review is unresolved, not a training winner. This is browser preparation, not local project or Portfolio II evidence. Implement convert_annotations in practice.py. Reviewed preference/model and fresh bounded-source helpers may be reused and disclosed, but implement your own declared mapping, accepted-ID policy, quarantine, original-label lineage and named losses. Return fresh data without mutating source; refuse unretainable inputs rather than invent a source report. Rerun every implemented practice group and keep the original CSV separate.
Change it, then build your own
One controlled change
Add new invented B and tie rows with a spaced/capitalized label, plus one new unknown label and extra field. Predict pair order, review verdict, original lineage and exact loss/reconciliation counts.
Your independent task
Implement convert_annotations in practice.py. Reviewed preference/model and fresh bounded-source helpers may be reused and disclosed, but implement your own declared mapping, accepted-ID policy, quarantine, original-label lineage and named losses. Return fresh data without mutating source; refuse unretainable inputs rather than invent a source report. Rerun every implemented practice group and keep the original CSV separate.
What success looks like
The build4 group checks A/B/tie semantics, source preservation, exact row reconciliation/loss counts, unknown/malformed/duplicate behavior, fresh ownership and whole-operation refusal for unretainable JSON data. The actual CSV demo agrees with those counts. This does not complete Portfolio II or establish automatic training eligibility.
Hint 1 · a question
Write the output and quarantine for A, B, tie, unknown and duplicate before coding. Which source facts must survive?
Hint 2 · a concept cue
Validate a candidate preference, reserve only accepted IDs, retain fresh refused sources and count extra keys only when conversion drops them.
Hint 3 · a localized example
For B, chosen_id comes from asset_b and rejected_id from asset_a. A lookup may use label.strip().casefold(), but source_label must retain the original string.
Need the complete worked solution?
Open dataset_tools/core.py and demo.py from the kit. Trace it, close it, then try fresh inputs in your own files. Treat the attempt as guided; seeing the solution does not award a practical pass.
Course help is guidance, not independent evidence. With JavaScript, opening help records guidance locally; otherwise note it in your README. Reset does not erase that history.
Repair a failed check
If B still chooses asset_a, trace the explicit swap. If unknown labels become winners, remove the default. If duplicate rows disappear, retain quarantine and physical row identity. If loss counts include quarantined extras, count only converted losses. If source labels lose case/spaces, separate lookup text from original lineage.
NotImplementedError means a practice stub is still unfinished. Read the failing test name and the last error line. Change one behavior, rerun that build, then rerun all implemented builds.
Show it works on new inputs
Supply a changed six-row source with A/B/tie/unknown/duplicate/invalid cases and one extra field. Show before/after samples, original lineage, named losses and input=converted+quarantined. Keep source bytes/manifest, explain review eligibility and disclose helper/reference assistance.
Self-review: name the input, result, refused case and reason. Your local test output and explanation are separate from a quiz score; this page does not certify a pass.
Keep the idea
A trustworthy conversion exposes lost and unresolved meaning. Useful counts never justify invented labels or silent data removal.
Module 6 checkpoint
Five questions, followed by the separate practical task above. JavaScript loads the scored questions; the build, files and hints remain available without it.