Does the rubric hold up under a second pass?
An anchored scale is a promise: that a 4 means the same thing every time someone writes one down. This page tests that promise on the CineQC rubric — the same clips scored twice by the same evaluator a week apart, then measured with quadratic-weighted Cohen's kappa per category. Where the two passes split, that category goes on a short list to investigate — the anchor wording is the first suspect, but not the only one.
Every number on this page is computed in your browser from the published dataset. Nothing is asserted that the dataset doesn't contain — and the dataset ships with the site so anyone can recompute it.
Reliability comes before accuracy
There is no ground truth for "is this AI clip good." Nobody can hand you the correct answer key, so accuracy is not measurable. What is measurable is whether the instrument is repeatable — whether two independent passes over the same clip land in the same place. An unreliable rubric cannot be accurate, because it doesn't produce a stable answer to be accurate about.
- Exact agreement — how often both passes chose the identical anchor. Intuitive, but flattering: on a 5-point scale, chance alone buys you some of it.
- Within ±1 — how often they landed adjacent. On an anchored scale, adjacent is usually a real match; the categories that fail even this are the broken ones.
- Mean absolute difference — the average gap in anchor points. Unlike the percentages, it doesn't hide the size of a disagreement.
- Quadratic-weighted Cohen's kappa (κw) — the headline. Corrects for chance agreement and punishes a 1-vs-5 split far harder than a 3-vs-4, which is exactly right for an ordinal scale. Read against the Landis & Koch bands: >0.8 almost perfect, 0.6–0.8 substantial, 0.4–0.6 moderate, below that the category needs investigating. Reported with a 95% bootstrap confidence interval, because at this sample size the interval is wide and the point estimate on its own would imply precision the data doesn't have.
- When kappa reads N/A. If both passes used a single value for every clip in a category, there is no variance for a chance correction to work on and kappa is not estimable — it is shown as N/A rather than as a perfect 1.0. Exact agreement still means something there and is still reported.
This is not a report card on the evaluator, and it is not proof about any single category either. It is a pilot: it finds where an anchored scale fails to reproduce itself, and turns that into a short list of categories worth investigating. Whether the cause is the wording, the workload, or the order the clips came in is the next study's job — one this design deliberately doesn't claim to settle.
Results
Loading dataset…
Agreement by category
Sorted widest-divergence first. The top row is where to look next, not a proven fault. Categories where kappa is not estimable sort to the bottom — a constant score is not a bad result.
The judgements that drove the gap
Largest splits first — these are the cases a rewrite has to make unambiguous.
Protocol
Written down before the data was collected, so the method can be checked against what was actually done.
- Sample, frozen before scoring. 25–30 clips are selected and locked first, against a written mix: generator, quality level, subject, amount of motion, and defect type. Freezing the set before any scoring is what stops the sample being quietly shaped by results. A set of only excellent or only broken clips manufactures agreement — everyone scores an obvious 5 the same way, and the kappa that results is meaningless.
- Independence. Each pass scores the clip without sight of the other pass.
- The Temporal Stability prefill is disabled for the study. CineQC normally seeds that one category from its automated measurement. Because the automated value is identical on both passes, an evaluator who accepts the default twice agrees with themselves by construction — the prefill manufactures agreement rather than measuring it, and on a seven-category rubric that one category would inflate every pooled figure on this page. It is therefore switched off during collection (
?noprefill=1), which shows a study-mode banner so a pass collected in the wrong mode is obvious rather than silently invalid. A dataset collected with the prefill on can still be salvaged: listing a category in the dataset'sexcludeFromHeadlinequarantines it — it is computed and shown in its own table below the results, but contributes nothing to the pooled figures. - Order, recorded. Pass 2 uses a randomized clip order, generated and written down before scoring starts, so fatigue and drift can't line up between passes — and so the ordering is auditable rather than "I shuffled it".
- No reconciliation. Passes are never discussed or adjusted before measurement. Reconciling first measures the conversation, not the rubric.
- Collection, exported every session. Each pass is exported as a CineQC share link and paired below. The working set lives in browser localStorage, which one cleared cache would destroy, so the dataset JSON is exported and saved at the end of every sitting rather than once at the end.
- Analyse before touching anything. The first read is on the original anchors, unchanged. Weak categories become hypotheses, not edits.
- Pre-registered read. Any category landing under κw = 0.6 is flagged for investigation. That threshold was fixed in advance of seeing results.
- Rewrite on a development subset, validate on a holdout. Anchor rewrites are developed against part of the set and tested on clips held back from that work — or by a third blinded pass. Same-clip before/after figures are reported as exploratory and labelled as such.
Add a clip to the set
Score a clip in CineQC, copy the share link, repeat for the second pass, then pair them here. The working set lives in this browser only until it's exported and committed.
Limitations, stated plainly
A reliability study that only reports its best number is advertising. These are the reasons to discount what's above.
- This is test–retest, not two people. Both passes are the same evaluator a week apart, so what it measures is intra-rater reliability — whether the anchors pull one person back to the same number. Treat it as an expected best case rather than a hard ceiling: two well-trained evaluators working from the same anchors can occasionally agree more consistently than one person does with themselves over time.
- Low agreement is a pointer, not a diagnosis. A weak category says the two passes diverged there. It does not establish why. Ambiguous anchor wording is the hypothesis this study is designed to generate, but fatigue, the order clips were seen in, something learned between passes, and ordinary evaluator drift all produce the same signature and are not separable here.
- Rewrite results on the same clips are exploratory. Re-scoring the clips that motivated a rewrite is biased toward showing improvement — regression to the mean alone will do some of it. A rewrite is only validated on clips held out from the diagnosis, or by a third blinded pass.
- Two passes, not a panel. Cohen's kappa describes these two passes. A production pipeline with many annotators wants Krippendorff's alpha across all of them; that needs raters this study doesn't have.
- Small n. Kappa on a modest clip set is indicative, not a confidence interval. Treat a single category's value as a pointer at the anchor text, not a precise measurement.
- Prompt Adherence is optional in the rubric, so its n is lower than the rest and its kappa is correspondingly noisier.
- Temporal Stability is prefilled from the automated pass for both evaluators — see the protocol. Its agreement is partly machine-supplied by construction.
- Reliability is not validity. Two passes agreeing proves the scale is repeatable. It does not prove the scale measures what matters to a viewer. That's a different study.
Measure the instrument, then trust it.
The rubric, the study, and the tool are the same artifact at different zoom levels. For evaluation partnerships, rubric design, or human-feedback pipeline work: marc.warfield@dynamicvibe.net