DVPPython Studio
Module 8: Model the evaluation domain / Build 3 of 4

Compose an explainable rubric

Apply interchangeable strategies to authored ratings with decimal weights, visible missing reviews and no automatic winner.

Runs on your computer · 70–100 minutes · no paid services

Download practice filesFiles, commands & notes

Without JavaScript, use the step links and keep your files on your computer.

One useful idea

A strategy packages one interchangeable calculation. Here each criterion exposes name, weight and evaluate(findings), returning CriterionResult with a Decimal rating or None and an explanation. Criterion is a structural Protocol: a trusted local object can satisfy the interface without inheriting a base class. MeanRating averages supplied authored ratings; LowestRating subclasses it but replaces the calculation. Neither observes a video or authenticates a reviewer.

For each asset and criterion, select only findings with that asset ID and criterion name, keeping finding IDs in the explanation. Mean motion ratings 5 and 3 produce 4; lighting rating 3 stays 3. With weights 0.6 and 0.4, the weighted mean is 3.6. Multiplying by 20 maps the classroom 1..5 range to 20..100, so the display is 72.00, not a calibrated quality probability.

Decimal("0.6") uses declared decimal text; Decimal(0.6) starts from a binary float approximation. Accept finite positive string/Decimal weights from 0.000001 through 1000, not floats or booleans. A fresh local decimal Context sets precision 28 and ROUND_HALF_UP, so unrelated caller precision, traps or exponent settings do not silently change the calculation. Quantize to 0.01 for the display while preserving the per-criterion explanation.

Missing authored findings yield None, not zero or a fabricated midpoint. If any criterion lacks a rating for an asset, that asset stays unscored with total None. Ratings for undeclared criteria remain visible in unused_finding_ids rather than disappearing. Reject duplicate criterion names, duplicate finding IDs and foreign membership. A strategy cannot invent a rating for an empty input.

The rubric accepts trusted local strategies, not downloaded plugins. It emits per-asset totals and supporting IDs, never a winner or preference. Changing weights or choosing lowest instead of mean is a disclosed policy change, not a factual discovery. Supplied model and strategy helpers are assistance; independently explain your aggregation and author a changed strategy for the transfer task.

Refresh first: Authored findings versus measurements, Classes, frozen values and methods.

Trace a finished example

import json
from pathlib import Path
from evalkit.model import EvaluationJob, ReviewerFinding

job_data = json.loads(Path("fixtures/job.json").read_text(encoding="utf-8"))
observations = json.loads(Path("fixtures/observations.json").read_text(encoding="utf-8"))
job = EvaluationJob.from_mapping(job_data)
findings = [ReviewerFinding.from_mapping(row) for row in observations["findings"]]

from evalkit.core import evaluate_rubric
from evalkit.rubric import MeanRating, LowestRating

mean = evaluate_rubric(job, findings, [MeanRating("motion", "0.6"), MeanRating("lighting", "0.4")])
lowest = evaluate_rubric(job, findings, [LowestRating("motion", "3"), LowestRating("lighting", "2")])
print(mean["assets"][0]["total"], lowest["assets"][0]["total"])
print(mean["assets"][1]["total"], mean["assets"][1]["status"])
print(mean["assets"][0]["criteria"][0]["finding_ids"])
motion_only = evaluate_rubric(job, findings, [MeanRating("motion", "1")])
print(motion_only["unused_finding_ids"])

Mean motion is 4; lowest motion is 3. Lighting is 3 for both, so the disclosed policies produce different totals. Asset b has no authored ratings and remains unscored. Excluding lighting retains its finding ID as unused evidence.

The finished implementation is in evalkit/core.py. Reading it is guided practice, not independent evidence.

Predict missing evidence

Should asset b receive 0.00 when no findings are supplied?

Compare your answer · self-reviewed

No. Zero invents a measured value outside this classroom rating scale. Missing review stays None/unscored and must not become a loss, tie or selected winner.

Find numeric drift

Why use Decimal("0.6") rather than Decimal(0.6)?

Compare your answer · self-reviewed

The string specifies a decimal value directly; the float carries a binary approximation. The rubric also owns its decimal context and rounding policy for reproducible calculation.

Recall strategy choice

Does changing mean to lowest reveal a new media fact?

Compare your answer · self-reviewed

No. It changes the aggregation policy over the same authored evidence. Disclose the strategy, weights, supporting IDs and missing reviews so another person can reproduce it.

Try the idea in this browser

Runs in this browser · optional preparation · local project checks remain separate

Try a small function before opening your local files. Python downloads when you choose Run; if it cannot load, your code stays here and the local kit still works. The worker executes on your device, not on a DVP server. Only run code you trust: this is not a hostile-code security sandbox.

JavaScript loads the practice controls. Python starts only after Run.

Read the browser task briefs without running Python

Guided compose an explainable rubric

Use invented in-memory values only. Reviewed scalar validators, immutable model classes and rubric strategies are supplied and disclosed. Implement your own evaluate_rubric, not the finished task function. These checks assess composition, not independent model/helper design, file publication, media quality or portfolio defense. Implement evaluate_rubric in practice.py with disclosed model and strategy helpers. Write your own validation, per-asset/per-criterion selection, finding-ID evidence, weighted Decimal aggregation and missing-review handling. Preserve undeclared finding IDs. Do not import the finished aggregator or turn totals into a winner.

Independent compose an explainable rubric

Use invented in-memory values only. Reviewed scalar validators, immutable model classes and rubric strategies are supplied and disclosed. Implement your own evaluate_rubric, not the finished task function. These checks assess composition, not independent model/helper design, file publication, media quality or portfolio defense. Implement evaluate_rubric in practice.py with disclosed model and strategy helpers. Write your own validation, per-asset/per-criterion selection, finding-ID evidence, weighted Decimal aggregation and missing-review handling. Preserve undeclared finding IDs. Do not import the finished aggregator or turn totals into a winner.

Change it, then build your own

One controlled change

Use motion weight 1 and lighting weight 1, then omit lighting from the declared rubric. Predict the new mean total and unused IDs. Remove one criterion’s findings and keep the asset unscored.

Your independent task

Implement evaluate_rubric in practice.py with disclosed model and strategy helpers. Write your own validation, per-asset/per-criterion selection, finding-ID evidence, weighted Decimal aggregation and missing-review handling. Preserve undeclared finding IDs. Do not import the finished aggregator or turn totals into a winner.

What success looks like

Build 3 checks mean/lowest/structural strategies, changed weights, supporting IDs, unused evidence, None/unscored, duplicate/foreign refusals and independence from caller decimal context. These checks establish an aggregation contract over supplied ratings, not an objective media ranking.

Hint 1 · a question

Trace motion values 5 and 3, lighting 3 and weights 0.6/0.4 by hand. Which result changes when the strategy becomes lowest?

Hint 2 · a concept cue

Filter by asset and criterion first. Keep supporting IDs, evaluate the selected tuple, and track missing results before deciding whether a total exists.

Hint 3 · a localized example

A complete weighted mean is sum(rating * weight) / sum(weight), then multiplied by 20 and quantized. An incomplete asset returns None instead of using a made-up zero.

Need the complete worked solution?

Open evalkit/core.py from the kit. Trace it, close it, then try fresh inputs in your own files. Treat the attempt as guided; seeing the solution does not award a practical pass.

Course help is guidance, not independent evidence. With JavaScript, opening help records guidance locally; otherwise note it in your README. Reset does not erase that history.

Repair a failed check

If an empty review gets a score, track completeness separately from the numeric accumulator. If results depend on caller precision, create the declared fresh Context. If findings disappear, retain unused IDs. If a machine row enters the rubric, require ReviewerFinding values rather than any dictionary with a number.

NotImplementedError means a practice stub is still unfinished. Read the failing test name and the last error line. Change one behavior, rerun that build, then rerun all implemented builds.

Show it works on new inputs

Author a trusted local highest-rating strategy, changed weights and a changed job with one unreviewed asset. Show totals, supporting/unused IDs and missing statuses. Explain the decimal scale and policy change without declaring an automatic winner; disclose reused model and strategy helpers.

Self-review: name the input, result, refused case and reason. Your local test output and explanation are separate from a quiz score; this page does not certify a pass.

Keep the idea

A useful rubric is reproducible and honest about absence. Its totals describe a declared policy over authored evidence, not an automatic verdict.