DVPPython Studio
Module 7: Extract signal from text / Build 4 of 4

Search a corpus with supporting evidence

Count accepted-note mentions and co-occurrences with reproducible spans, honest refusals and human-review limits.

Runs on your computer · 70–100 minutes · no paid services

Download practice filesFiles, commands & notes

Without JavaScript, use the step links and keep your files on your computer.

One useful idea

A corpus is a collection of notes. This search index helps a person find mentions; it does not inspect video or decide whether a reviewer is correct. Accept an iterable of exact note_id/note dictionaries, not a bare string/bytes/dict. IDs are opaque 1–64 ASCII letters/digits/underscore/hyphen; keep secrets out of IDs because that grammar is not redaction. Invalid rows use schema quarantine. The first accepted ID wins; later valid duplicates use duplicate. A prior refused row does not reserve an ID.

Normalize each accepted note, mask declared leakage shapes, then search that redacted normalized text. The fixed English lexicon is continuity: mismatch/wardrobe change; lighting: flicker/underexposed; motion: jitter/stutter. Match ASCII English case variants with Unicode word boundaries; jittery, jitteré and stutter_2 are not those words. Normalization makes wardrobe followed by a newline then change into the declared phrase. Bracket tags alone do not imply lexicon matches.

Keep every tag/start/exclusive-end observation sorted by position. Its basis is now NORMALIZED redacted text, not original text. Count each tag once per accepted note, even when it appears three times. itertools.combinations(tags, 2) enumerates unique unordered tag pairs; count a pair once per note as well. Return all three zero-inclusive category counts, all sorted pair counts and ordered evidence with row/note_id/tags/redacted_note/matches. A set provides uniqueness; Counter.update adds one count per item supplied.

Reconcile input_rows=accepted_rows+quarantined_rows and accepted_rows=matched_rows+unmatched_rows. Diagnostic quarantine retains row, valid supplied ID or None and controlled code, not raw private note/extra fields; it is not a complete backup. Keep source separately. At most 10,000 physical rows and 1,000,000 supplied code points across valid rows including duplicates; exceeding a budget refuses the whole operation. Per-note text limits still apply. IDs/evidence grow in memory; unexpected iterator faults propagate rather than become fake partial success.

“No jitter” still contains the lexicon word. Negation, context, sarcasm and reviewer accuracy remain unresolved. These are mentions, not verified defects or a failure rate. Optional demo export writes a NEW JSON file, refuses existing/known-link/reserved/parent-component paths and verifies actual bytes/hash/typed read-back. Partial new files can remain after failure, with no success receipt or automatic rollback. Serialization parity is not analysis correctness, privacy, independent work or completed Portfolio II.

Refresh first: Fixed-shape masking, Declared normalization, Accepted IDs and row reconciliation.

Trace a finished example

from demo import read_fixture
from text_tools.core import search_corpus

report = search_corpus(read_fixture("corpus"))
print(report["input_rows"], report["accepted_rows"], report["quarantined_rows"])
print(report["matched_rows"], report["unmatched_rows"])
print(report["counts"])
print([item["note_id"] for item in report["evidence"]])
print(report["evidence"][0]["redacted_note"])

Six actual fixture rows produce four accepted notes and two diagnostic refusals. Two accepted notes support the category counts; two have no lexicon mention after masking. Duplicate stutter is excluded. No jitter still supplies a motion mention and requires human interpretation, not a defect verdict.

The finished implementation is in text_tools/core.py and demo.py. Reading it is guided practice, not independent evidence.

Predict repetition

One accepted note says jitter three times. Is motion counted three times?

Compare your answer · self-reviewed

No. Keep three spans but count one accepted note with the category. Occurrences and per-note counts answer different questions.

Find false judgment

Can a motion count from “No jitter” be called a verified motion failure?

Compare your answer · self-reviewed

No. This lexicon index does not resolve negation or inspect media. Use its supporting text for human review and label the count mentions.

Recall accepted identity

Does an invalid row consume an ID that a later valid row may use?

Compare your answer · self-reviewed

No. Reserve only accepted IDs. Preserve the invalid row’s diagnostic separately, then accept the first valid record for that ID.

Try the idea in this browser

Runs in this browser · optional preparation · local project checks remain separate

Try a small function before opening your local files. Python downloads when you choose Run; if it cannot load, your code stays here and the local kit still works. The worker executes on your device, not on a DVP server. Only run code you trust: this is not a hostile-code security sandbox.

JavaScript loads the practice controls. Python starts only after Run.

Read the browser task briefs without running Python

Guided search a corpus with supporting evidence

Use invented in-memory text only. Reviewed bounded-input policies, fixed patterns and vocabulary are supplied and disclosed; reviewed normalization/masking helpers are also supplied for integration. Write your own search_corpus; the finished task function is not supplied. No file/export/provider operation runs and these checks are not local project or Portfolio II evidence. Implement search_corpus in practice.py with exact row validation, accepted-ID policy, bounded totals, reviewed normalization/masking, fixed boundary-aware search, per-note counts and actual supporting spans. Helper use must be disclosed; reusing reviewed normalize_note/scan_leakage is not independent implementation of those tasks. Do not import the finished search function or swallow unexpected source-iterator failures.

Independent search a corpus with supporting evidence

Use invented in-memory text only. Reviewed bounded-input policies, fixed patterns and vocabulary are supplied and disclosed; reviewed normalization/masking helpers are also supplied for integration. Write your own search_corpus; the finished task function is not supplied. No file/export/provider operation runs and these checks are not local project or Portfolio II evidence. Implement search_corpus in practice.py with exact row validation, accepted-ID policy, bounded totals, reviewed normalization/masking, fixed boundary-aware search, per-note counts and actual supporting spans. Helper use must be disclosed; reusing reviewed normalize_note/scan_leakage is not independent implementation of those tasks. Do not import the finished search function or swallow unexpected source-iterator failures.

Change it, then build your own

One controlled change

Add a new note repeating two categories, one negated mention, one duplicate and one malformed row. Predict accepted/matched counts and one co-occurrence pair; reproduce them from evidence, not from the raw number of word appearances.

Your independent task

Implement search_corpus in practice.py with exact row validation, accepted-ID policy, bounded totals, reviewed normalization/masking, fixed boundary-aware search, per-note counts and actual supporting spans. Helper use must be disclosed; reusing reviewed normalize_note/scan_leakage is not independent implementation of those tasks. Do not import the finished search function or swallow unexpected source-iterator failures.

What success looks like

The build4 tests check changed Unicode/case boundaries, per-note rather than occurrence counts, normalized offset basis, redaction before search, schema/duplicate reconciliation, empty evidence, inclusive budgets and visible iterator faults. NFC can expand text: 6,001 copies of U+0344 become 12,002 code points. If normalized text exceeds the downstream 12,000-code-point scanner budget, refuse the WHOLE operation with normalized note budget exceeded, not truncation, quarantine or partial success. The optional local export is a real NEW JSON file, not a browser filesystem claim or independent portfolio certificate.

Hint 1 · a question

Write one evidence entry with its exact redacted normalized text and spans. Derive its tag set before counting anything.

Hint 2 · a concept cue

Separate row validation, accepted identity, normalization/masking, observations and aggregation. Count unique tags/pairs per accepted note and retain controlled refusal diagnostics.

Hint 3 · a localized example

combinations(["lighting", "motion"], 2) yields one pair. Repetition in the same note adds supporting spans, not another note to that pair count.

Need the complete worked solution?

Open text_tools/core.py and demo.py from the kit. Trace it, close it, then try fresh inputs in your own files. Treat the attempt as guided; seeing the solution does not award a practical pass.

Course help is guidance, not independent evidence. With JavaScript, opening help records guidance locally; otherwise note it in your README. Reset does not erase that history.

Repair a failed check

If counts exceed evidence note sets, deduplicate tags before updating totals and pairs. If an email contributes jitter, mask before searching. If invalid rows reserve IDs, move reservation after validation. If a later source fault prints success, let it propagate. If negated mentions become defects, correct the interpretation label rather than pretending regex understands judgment.

NotImplementedError means a practice stub is still unfinished. Read the failing test name and the last error line. Change one behavior, rerun that build, then rerun all implemented builds.

Show it works on new inputs

Author a changed six-row corpus with repetition, negation, two categories, duplicate, invalid row and Unicode text. Reconcile both equations and reproduce category/pair counts from evidence. Run your own optional export into a NEW filename, retain original source separately and explain privacy/mention/helper limits.

Self-review: name the input, result, refused case and reason. Your local test output and explanation are separate from a quiz score; this page does not certify a pass.

Keep the idea

Counts are useful when a person can reproduce and interpret them. Search evidence is not automated media judgment or completed Portfolio II.

Module 7 checkpoint

Five questions, followed by the separate practical task above. JavaScript loads the scored questions; the build, files and hints remain available without it.