Methodology
The scoring, blinding, confidence, export, and reliability rules that make EvalForge evaluations reproducible.
Scoring: the 1–5 mapping
Each rubric category is scored on an anchored 1–5 scale, where every level names an operational state rather than a bare adjective. A candidate's normalized score is the weighted mean of the categories it actually completed, multiplied by 20, giving a 0–100 scale (a rating of 5 across all categories is 100).
Optional-category rule: required categories must be scored on every candidate in a group. An optional category may be skipped — but if it is scored for any candidate in a group, it must be scored for all of them, so comparisons stay symmetric.
Confidence boundaries
Confidence is classified by the top-two margin on the 0–100 scale: a margin of 15 or more is high, 5 to just under 15 is medium, and below 5 is low. Equal rounded top scores are flagged as a calculated tie.
Tie and override rules
The calculated winner is a suggestion, never the final word. Confirming accepts it. An override records a different winner and requires a written rationale; it never rewrites the category scores. A human tie and an exclusion also require a rationale. Every decision keeps both the calculated order and the human outcome.
Blinding and identity separation
At import, groups and candidates are shuffled by a seeded, reproducible order and relabeled with neutral codes (G-####/C-##). The evaluator-facing manifest carries codes only — never a filename, model, version, or seed. Real identities live in a separate identity map that is reattached only through a deliberate, audited unblinding.
Export eligibility
Preference records export as CSV or JSONL (one object per line). Confirmed and overridden preferences are eligible; low-confidence and calculated-tie preferences export but are flagged. Human ties, exclusions, invalid groups, and undecided groups are ineligible and are held back from the dataset.
Reliability and its limitations
Rubric reliability is reported with weighted kappa for category ratings and preference agreement across evaluators, with bootstrap intervals. Kappa has known limitations: it is sensitive to category prevalence and marginal distributions, so a high raw-agreement rate can still yield a modest kappa. We report both, and we do not publish any reliability number until a second evaluator's independent subset exists.