DVPPython Studio
Module 7: Extract signal from text / Build 3 of 4

Normalize notes without rewriting meaning

Create a stable lookup form while preserving original judgment, Unicode distinctions and unfamiliar tags.

Runs on your computer · 50–75 minutes · no paid services

Download practice filesFiles, commands & notes

Without JavaScript, use the step links and keep your files on your computer.

One useful idea

Normalization creates a comparison form, not a new reviewer judgment. Keep original first. unicodedata.normalize("NFC", original) composes canonically equivalent Unicode sequences; for example an e plus combining accent can become é. NFC is different from NFKC compatibility normalization: fullwidth letters and ligatures must remain distinct in this contract. No ASCII folding or lowercasing of the whole note is allowed.

Then use " ".join(composed.split()) to collapse whitespace into single spaces and trim its edges. Compare each stage separately: record unicode_nfc only if NFC changed text and whitespace only if the second stage changed it. Return ordered changes, not an unconditional “cleaned” badge. Preserve case, punctuation, negation and zero-width format characters. A zero-width character is not removed merely because it looks like spacing.

Observe bracket tags with an ASCII letter followed by up to 31 ASCII letters/digits/underscore/hyphen. Regex finditer plus group(1) reads the tag content; casefold only that valid ASCII label for lookup. Known motion/lighting/continuity/composition labels enter sorted unique tags; other matching labels enter unknown_tags. Keep all bracket text in normalized. Malformed brackets remain text rather than invented tags. An explicit bracket scanner can implement the same policy without regex.

Return normalizer_version=1, original, normalized, tags, unknown_tags and changes, with fresh lists. The shared 12,000-code-point input/control policy applies; empty notes are valid observations. Original retention deliberately keeps potentially private source data. This function is not redaction, a share-safe log or proof that a finding is accurate; nothing sends it to the tutor automatically. “No jitter!” must retain its negation and punctuation.

Refresh first: Original text and offset boundaries, Split text into fields.

Trace a finished example

from text_tools.core import normalize_note

source = "  Cafe\u0301\u00a0海\n[MOTION] [Tone]\t"
report = normalize_note(source)
print(report["normalized"])
print(report["tags"], report["unknown_tags"])
print(report["changes"])
print(report["original"] == source)
print(normalize_note("No jitter! A ff")["normalized"])

NFC composes the decomposed accent. Whitespace collapse joins the line without erasing tag text. MOTION is known; Tone remains unknown instead of disappearing. The original stays exact, and the second note keeps negation, fullwidth A and the ligature.

The finished implementation is in text_tools/core.py. Reading it is guided practice, not independent evidence.

Predict compatibility

Should NFC turn A into A in this contract?

Compare your answer · self-reviewed

No. That is a compatibility distinction. NFC preserves it; silently choosing NFKC or ASCII folding would change the declared policy.

Find missing evidence

Why is deleting [Tone] a bad way to handle an unknown tag?

Compare your answer · self-reviewed

It removes supplied meaning. Preserve bracket text and record a valid unfamiliar tag in unknown_tags, rather than invent a canonical category or discard it.

Recall stage accounting

If NFC does nothing but whitespace changes, which change should be recorded?

Compare your answer · self-reviewed

Only whitespace. Compare original to the NFC stage and that stage to the collapsed text separately; a generic changed flag cannot explain which transformation happened.

Try the idea in this browser

Runs in this browser · optional preparation · local project checks remain separate

Try a small function before opening your local files. Python downloads when you choose Run; if it cannot load, your code stays here and the local kit still works. The worker executes on your device, not on a DVP server. Only run code you trust: this is not a hostile-code security sandbox.

JavaScript loads the practice controls. Python starts only after Run.

Read the browser task briefs without running Python

Guided normalize notes without rewriting meaning

Use invented in-memory text only. Reviewed bounded-input policies, fixed patterns and vocabulary are supplied and disclosed. Write your own normalize_note; the finished task function is not supplied. No file/export/provider operation runs and these checks are not local project or Portfolio II evidence. Implement normalize_note in practice.py. You may disclose the reviewed input policy, tag pattern and allowed vocabulary, but write your own NFC/whitespace stages, exact original retention, tag classification and fresh result lists. Do not import the finished normalizer. Neither regex nor a manual bracket parser may rewrite the whole judgment.

Independent normalize notes without rewriting meaning

Use invented in-memory text only. Reviewed bounded-input policies, fixed patterns and vocabulary are supplied and disclosed. Write your own normalize_note; the finished task function is not supplied. No file/export/provider operation runs and these checks are not local project or Portfolio II evidence. Implement normalize_note in practice.py. You may disclose the reviewed input policy, tag pattern and allowed vocabulary, but write your own NFC/whitespace stages, exact original retention, tag classification and fresh result lists. Do not import the finished normalizer. Neither regex nor a manual bracket parser may rewrite the whole judgment.

Change it, then build your own

One controlled change

Use already composed Café with no extra spaces, then decomposed text with a newline, unknown [New_Tag] and a zero-width character. Predict each stage’s change list and which original distinctions survive.

Your independent task

Implement normalize_note in practice.py. You may disclose the reviewed input policy, tag pattern and allowed vocabulary, but write your own NFC/whitespace stages, exact original retention, tag classification and fresh result lists. Do not import the finished normalizer. Neither regex nor a manual bracket parser may rewrite the whole judgment.

What success looks like

The build3 tests compare exact original and normalized forms, meaningful Unicode/whitespace changes, known/unknown/malformed tags, punctuation/negation/compatibility/zero-width retention and ownership. Passing them establishes this formatting policy, not that the note is accurate or safe to upload.

Hint 1 · a question

Which exact data must survive before normalization begins? Write expected original, NFC stage and whitespace stage side by side.

Hint 2 · a concept cue

Record a change only when that stage differs. Build a unique tag set from valid bracket observations, then separate known from unknown without deleting text.

Hint 3 · a localized example

" ".join(text.split()) collapses whitespace. It is not redaction or a reason to remove punctuation, “No”, or an unfamiliar bracket label.

Need the complete worked solution?

Open text_tools/core.py from the kit. Trace it, close it, then try fresh inputs in your own files. Treat the attempt as guided; seeing the solution does not award a practical pass.

Course help is guidance, not independent evidence. With JavaScript, opening help records guidance locally; otherwise note it in your README. Reset does not erase that history.

Repair a failed check

If original equals the cleaned string instead of supplied text, save it before any transformation. If fullwidth or ligature characters change, inspect the NFC/NFKC argument. If negation vanishes, remove judgment rewriting. If unknown tags disappear, preserve text and classify rather than filter it out.

NotImplementedError means a practice stub is still unfinished. Read the failing test name and the last error line. Change one behavior, rerun that build, then rerun all implemented builds.

Show it works on new inputs

Author changed Unicode notes with canonical accent composition, whitespace, negation, a compatibility character and an unknown tag. Show originals beside normalized forms and ordered changes. Explain what lookup normalization loses and why original retention requires privacy care.

Self-review: name the input, result, refused case and reason. Your local test output and explanation are separate from a quiz score; this page does not certify a pass.

Keep the idea

Preserve original judgment and explain every transformation. A stable lookup form is not a rewritten verdict.