DVPPython Studio
Module 8: Model the evaluation domain / Build 2 of 4

Keep machine signals separate from findings

Preserve two observation layers, check their membership and avoid turning a detector measurement into a human verdict.

Runs on your computer · 60–90 minutes · no paid services

Download practice filesFiles, commands & notes

Without JavaScript, use the step links and keep your files on your computer.

One useful idea

A machine measurement and a reviewer finding answer different questions. MachineSignal records a supplied metric, finite value in 0..1, producer and limitations. ReviewerFinding records a supplied criterion, authored integer rating in 1..5, author ID and evidence text. The fixture signal 0.99 is invented. It is not a probability calibrated against real media, a quality score or permission to choose a winner.

Validate every row against its exact schema before collecting it. Check that job_id matches this job and asset_id belongs to one of its two assets. Signal IDs must be unique within signals; finding IDs must be unique within findings. An empty layer is valid. Inputs are lists with at most 1,000 observations each; the limit is a declared resource budget, not a measure of completeness.

observation_bundle constructs a typed job, then traverses each layer using its own class and identity field. A set detects duplicates; fresh serialized dictionaries preserve values without retaining caller-owned rows. A machine row cannot masquerade as a finding by changing a label. The required fields and numeric policies differ. Invalid membership or schema refuses the whole operation instead of producing a partial success report.

The result keeps signals and findings in separate lists. review_status is unreviewed when findings are empty and findings_present otherwise. That second label means supplied findings exist, not that their author is authenticated or their evidence has been independently verified. No chosen, winner or verdict field is produced. Contradictory supplied observations stay visible for a person to examine.

The browser task uses invented in-memory rows only. The local kit exercises the same boundary without opening media. Use no private prompt or reviewer evidence in course notes or the tutor unless you have reviewed what you intend to share. Reviewed model classes and the job serializer are supplied assistance; your own collection and membership logic remain the task.

Refresh first: Typed job ownership, Mentions are not verified defects.

Trace a finished example

import json
from pathlib import Path
from evalkit.core import observation_bundle

job = json.loads(Path("fixtures/job.json").read_text(encoding="utf-8"))
rows = json.loads(Path("fixtures/observations.json").read_text(encoding="utf-8"))
rows["findings"][0]["rating"] = 1
report = observation_bundle(job, rows["signals"], rows["findings"])
print(report["signals"][0]["value"], report["findings"][0]["rating"], report["review_status"])
print(observation_bundle(job, rows["signals"], [])["review_status"])
print("winner" in report, "verdict" in report)
try:
    observation_bundle(job, [], rows["signals"])
except ValueError:
    print("signal as finding refused")

The supplied measurement stays 0.99 even when the authored rating changes to 1. Neither is rewritten to agree. No findings means unreviewed; a signal cannot supply the missing human layer or create a verdict.

The finished implementation is in evalkit/core.py. Reading it is guided practice, not independent evidence.

Predict disagreement

Should an authored rating of 1 be changed to 5 because the signal is 0.99?

Compare your answer · self-reviewed

No. Preserve the supplied measurement and authored finding in their separate layers. Their disagreement requires interpretation, not automatic rewriting or winner selection.

Find the membership bug

A finding has the right asset ID but a different job ID. Is it valid here?

Compare your answer · self-reviewed

No. Both membership checks matter. Refuse a foreign observation instead of attaching its evidence to this job merely because an asset label matches.

Recall status meaning

Does findings_present verify the named reviewer?

Compare your answer · self-reviewed

No. It states that supplied findings exist. IDs and evidence are supplied claims; this model does not authenticate authorship or independently verify judgment.

Try the idea in this browser

Runs in this browser · optional preparation · local project checks remain separate

Try a small function before opening your local files. Python downloads when you choose Run; if it cannot load, your code stays here and the local kit still works. The worker executes on your device, not on a DVP server. Only run code you trust: this is not a hostile-code security sandbox.

JavaScript loads the practice controls. Python starts only after Run.

Read the browser task briefs without running Python

Guided keep machine signals separate from findings

Use invented in-memory values only. Reviewed scalar validators, immutable model classes and rubric strategies are supplied and disclosed; the job serializer is supplied for composition. Implement your own observation_bundle, not the finished task function. These checks assess composition, not independent model/helper design, file publication, media quality or portfolio defense. Implement observation_bundle in practice.py. Reuse the disclosed typed models but write your own bounded layer traversal, per-layer duplicate checks, job/asset membership and fresh serialization. Preserve both layers without generating a winner. Do not import the finished core function or silently drop invalid observations.

Independent keep machine signals separate from findings

Use invented in-memory values only. Reviewed scalar validators, immutable model classes and rubric strategies are supplied and disclosed; the job serializer is supplied for composition. Implement your own observation_bundle, not the finished task function. These checks assess composition, not independent model/helper design, file publication, media quality or portfolio defense. Implement observation_bundle in practice.py. Reuse the disclosed typed models but write your own bounded layer traversal, per-layer duplicate checks, job/asset membership and fresh serialization. Preserve both layers without generating a winner. Do not import the finished core function or silently drop invalid observations.

Change it, then build your own

One controlled change

Remove all findings while keeping the signal, then add a changed valid finding. Try a duplicate finding ID and a foreign job ID. Predict status and refusal without treating the measurement as a review.

Your independent task

Implement observation_bundle in practice.py. Reuse the disclosed typed models but write your own bounded layer traversal, per-layer duplicate checks, job/asset membership and fresh serialization. Preserve both layers without generating a winner. Do not import the finished core function or silently drop invalid observations.

What success looks like

Build 2 checks actual changed observations, duplicate and foreign identities, boolean/range/type refusals, empty-layer status and fresh ownership. A pass verifies these collection contracts, not media quality, authenticated review or a completed evaluation.

Hint 1 · a question

Write one measurement and one finding with all declared fields. Which fields differ, and which IDs connect each to the job?

Hint 2 · a concept cue

Validate before collecting. A separate seen set for each layer detects duplicate identities without merging measurements and authored judgments.

Hint 3 · a localized example

The status is based on whether the validated finding layer is empty. It is not based on signal magnitude, quality or reviewer authentication.

Need the complete worked solution?

Open evalkit/core.py from the kit. Trace it, close it, then try fresh inputs in your own files. Treat the attempt as guided; seeing the solution does not award a practical pass.

Course help is guidance, not independent evidence. With JavaScript, opening help records guidance locally; otherwise note it in your README. Reset does not erase that history.

Repair a failed check

If a high measurement creates findings, remove the invented conversion. If a foreign row passes, check both IDs. If duplicates disappear silently, refuse rather than deduplicate evidence. If a report changes after input mutation, serialize fresh values. Keep unexpected faults visible instead of reporting success.

NotImplementedError means a practice stub is still unfinished. Read the failing test name and the last error line. Change one behavior, rerun that build, then rerun all implemented builds.

Show it works on new inputs

Create two changed assets, conflicting machine/human observations, an empty finding layer, duplicate IDs and foreign membership. Save actual results/refusals and explain why the layers must remain distinct. Disclose model helpers and do not describe invented fixture values as measured media evidence.

Self-review: name the input, result, refused case and reason. Your local test output and explanation are separate from a quiz score; this page does not certify a pass.

Keep the idea

Observation, authored judgment and verified outcome are different states. Keeping them separate makes disagreement and missing evidence visible.