Synthetic DIY-repair Q&A + data-quality gate + human-calibrated LLM-as-judge eval + prompt correction. newline Miniproject 1.
-
Updated
Jul 3, 2026 - Python
Synthetic DIY-repair Q&A + data-quality gate + human-calibrated LLM-as-judge eval + prompt correction. newline Miniproject 1.
An eval harness that gates deployment and measures its own judges. Three LLM judges score a system under test against a rubric, plus two non-voting shadows, and the gate fails CI on a golden-set regression. Reports human agreement, Cohen's kappa, self-consistency, judge error correlation, and refuses a gate inside the judges' own noise floor.
Add a description, image, and links to the judge-calibration topic page so that developers can more easily learn about it.
To associate your repository with the judge-calibration topic, visit your repo's landing page and select "manage topics."