One useful idea
A strategy packages one interchangeable calculation. Here each criterion exposes name, weight and evaluate(findings), returning CriterionResult with a Decimal rating or None and an explanation. Criterion is a structural Protocol: a trusted local object can satisfy the interface without inheriting a base class. MeanRating averages supplied authored ratings; LowestRating subclasses it but replaces the calculation. Neither observes a video or authenticates a reviewer.
For each asset and criterion, select only findings with that asset ID and criterion name, keeping finding IDs in the explanation. Mean motion ratings 5 and 3 produce 4; lighting rating 3 stays 3. With weights 0.6 and 0.4, the weighted mean is 3.6. Multiplying by 20 maps the classroom 1..5 range to 20..100, so the display is 72.00, not a calibrated quality probability.
Decimal("0.6") uses declared decimal text; Decimal(0.6) starts from a binary float approximation. Accept finite positive string/Decimal weights from 0.000001 through 1000, not floats or booleans. A fresh local decimal Context sets precision 28 and ROUND_HALF_UP, so unrelated caller precision, traps or exponent settings do not silently change the calculation. Quantize to 0.01 for the display while preserving the per-criterion explanation.
Missing authored findings yield None, not zero or a fabricated midpoint. If any criterion lacks a rating for an asset, that asset stays unscored with total None. Ratings for undeclared criteria remain visible in unused_finding_ids rather than disappearing. Reject duplicate criterion names, duplicate finding IDs and foreign membership. A strategy cannot invent a rating for an empty input.
The rubric accepts trusted local strategies, not downloaded plugins. It emits per-asset totals and supporting IDs, never a winner or preference. Changing weights or choosing lowest instead of mean is a disclosed policy change, not a factual discovery. Supplied model and strategy helpers are assistance; independently explain your aggregation and author a changed strategy for the transfer task.
Refresh first: Authored findings versus measurements, Classes, frozen values and methods.
Trace a finished example
import json
from pathlib import Path
from evalkit.model import EvaluationJob, ReviewerFinding
job_data = json.loads(Path("fixtures/job.json").read_text(encoding="utf-8"))
observations = json.loads(Path("fixtures/observations.json").read_text(encoding="utf-8"))
job = EvaluationJob.from_mapping(job_data)
findings = [ReviewerFinding.from_mapping(row) for row in observations["findings"]]
from evalkit.core import evaluate_rubric
from evalkit.rubric import MeanRating, LowestRating
mean = evaluate_rubric(job, findings, [MeanRating("motion", "0.6"), MeanRating("lighting", "0.4")])
lowest = evaluate_rubric(job, findings, [LowestRating("motion", "3"), LowestRating("lighting", "2")])
print(mean["assets"][0]["total"], lowest["assets"][0]["total"])
print(mean["assets"][1]["total"], mean["assets"][1]["status"])
print(mean["assets"][0]["criteria"][0]["finding_ids"])
motion_only = evaluate_rubric(job, findings, [MeanRating("motion", "1")])
print(motion_only["unused_finding_ids"])Mean motion is 4; lowest motion is 3. Lighting is 3 for both, so the disclosed policies produce different totals. Asset b has no authored ratings and remains unscored. Excluding lighting retains its finding ID as unused evidence.
The finished implementation is in evalkit/core.py. Reading it is guided practice, not independent evidence.
Predict missing evidence
Should asset b receive 0.00 when no findings are supplied?
Compare your answer · self-reviewed
No. Zero invents a measured value outside this classroom rating scale. Missing review stays None/unscored and must not become a loss, tie or selected winner.
Find numeric drift
Why use Decimal("0.6") rather than Decimal(0.6)?
Compare your answer · self-reviewed
The string specifies a decimal value directly; the float carries a binary approximation. The rubric also owns its decimal context and rounding policy for reproducible calculation.
Recall strategy choice
Does changing mean to lowest reveal a new media fact?
Compare your answer · self-reviewed
No. It changes the aggregation policy over the same authored evidence. Disclose the strategy, weights, supporting IDs and missing reviews so another person can reproduce it.
Try the idea in this browser
Runs in this browser · optional preparation · local project checks remain separate
Try a small function before opening your local files. Python downloads when you choose Run; if it cannot load, your code stays here and the local kit still works. The worker executes on your device, not on a DVP server. Only run code you trust: this is not a hostile-code security sandbox.
JavaScript loads the practice controls. Python starts only after Run.
The tutor button only prepares a question locally. Review it and choose Send yourself; no code is sent merely by running or opening a lesson.
Output
Errors and check feedback
Read the browser task briefs without running Python
Guided compose an explainable rubric
Use invented in-memory values only. Reviewed scalar validators, immutable model classes and rubric strategies are supplied and disclosed. Implement your own evaluate_rubric, not the finished task function. These checks assess composition, not independent model/helper design, file publication, media quality or portfolio defense. Implement evaluate_rubric in practice.py with disclosed model and strategy helpers. Write your own validation, per-asset/per-criterion selection, finding-ID evidence, weighted Decimal aggregation and missing-review handling. Preserve undeclared finding IDs. Do not import the finished aggregator or turn totals into a winner.
Independent compose an explainable rubric
Use invented in-memory values only. Reviewed scalar validators, immutable model classes and rubric strategies are supplied and disclosed. Implement your own evaluate_rubric, not the finished task function. These checks assess composition, not independent model/helper design, file publication, media quality or portfolio defense. Implement evaluate_rubric in practice.py with disclosed model and strategy helpers. Write your own validation, per-asset/per-criterion selection, finding-ID evidence, weighted Decimal aggregation and missing-review handling. Preserve undeclared finding IDs. Do not import the finished aggregator or turn totals into a winner.
Change it, then build your own
One controlled change
Use motion weight 1 and lighting weight 1, then omit lighting from the declared rubric. Predict the new mean total and unused IDs. Remove one criterion’s findings and keep the asset unscored.
Your independent task
Implement evaluate_rubric in practice.py with disclosed model and strategy helpers. Write your own validation, per-asset/per-criterion selection, finding-ID evidence, weighted Decimal aggregation and missing-review handling. Preserve undeclared finding IDs. Do not import the finished aggregator or turn totals into a winner.
What success looks like
Build 3 checks mean/lowest/structural strategies, changed weights, supporting IDs, unused evidence, None/unscored, duplicate/foreign refusals and independence from caller decimal context. These checks establish an aggregation contract over supplied ratings, not an objective media ranking.
Hint 1 · a question
Trace motion values 5 and 3, lighting 3 and weights 0.6/0.4 by hand. Which result changes when the strategy becomes lowest?
Hint 2 · a concept cue
Filter by asset and criterion first. Keep supporting IDs, evaluate the selected tuple, and track missing results before deciding whether a total exists.
Hint 3 · a localized example
A complete weighted mean is sum(rating * weight) / sum(weight), then multiplied by 20 and quantized. An incomplete asset returns None instead of using a made-up zero.
Need the complete worked solution?
Open evalkit/core.py from the kit. Trace it, close it, then try fresh inputs in your own files. Treat the attempt as guided; seeing the solution does not award a practical pass.
Course help is guidance, not independent evidence. With JavaScript, opening help records guidance locally; otherwise note it in your README. Reset does not erase that history.
Repair a failed check
If an empty review gets a score, track completeness separately from the numeric accumulator. If results depend on caller precision, create the declared fresh Context. If findings disappear, retain unused IDs. If a machine row enters the rubric, require ReviewerFinding values rather than any dictionary with a number.
NotImplementedError means a practice stub is still unfinished. Read the failing test name and the last error line. Change one behavior, rerun that build, then rerun all implemented builds.
Show it works on new inputs
Author a trusted local highest-rating strategy, changed weights and a changed job with one unreviewed asset. Show totals, supporting/unused IDs and missing statuses. Explain the decimal scale and policy change without declaring an automatic winner; disclose reused model and strategy helpers.
Self-review: name the input, result, refused case and reason. Your local test output and explanation are separate from a quiz score; this page does not certify a pass.
Keep the idea
A useful rubric is reproducible and honest about absence. Its totals describe a declared policy over authored evidence, not an automatic verdict.