ICC: Grading the Graders
What is the intraclass correlation coefficient (ICC) and why does it matter? When AI grades AI, someone needs to grade the grading. Here’s the measure we use and why that number alone isn’t the whole story.
Everyone wants a single number. How safe is this AI? How fair? Give me a score and let me continue. We feel the pull—we publish scores too. But a score is only worth something if it is reliable. If you run the same test tomorrow or hand it to a different evaluator, and a trustworthy score should barely change. So before we ask what an AI scores, we should ask: how do we know the score itself is any good?
The AI Ethics Index uses a panel of AI models to score the evidence for each indicator. Three judges drawn from three different model families score the same evidence independently. We do this on purpose. No single model gets to decide a result on its own, so no single model’s blind spots determine the verdict. And the panel is not the last step: a human evaluator also reviews the panel’s scoring before publication.
This raises an immediate question: do the AI judges agree? If the judges disagree substantially, their average may conceal an unstable measurement. To address this, we measure whether they agree, and we treat that agreement as a number to be analyzed in its own right.
How do we check scores?
We use a statistical tool called the intraclass correlation coefficient (ICC). ICC compares variation among the cases being scored with variation attributable to disagreement among the judges. A high ICC indicates that judge disagreement is small relative to the differences among the cases. A low ICC indicates that judge-to-judge variation is large enough to limit confidence in the resulting score.
There’s one important choice we make with this tool: we use the “absolute agreement” version instead of the looser “consistency” version. The consistency version only asks whether judges rank cases similarly. Under this criterion, a systematic difference in level—for example, one judge assigning every case a lower score—may still produce high consistency.
The absolute-agreement standard checks whether judges land on the same number, and penalizes systematic differences in scoring. That distinction matters because we use these scores against fixed thresholds and in audits, where the exact value matters, not just the ranking.
Since our judges are different AI models, each with its own built-in grading tendencies, we need our checking method to catch, rather than smooth over, systematic severity or leniency. The absolute-agreement standard allows us to detect these differences, providing greater insight into how well our AI judges truly agree.
Two numbers, not one
There’s a subtlety worth being upfront about. The strictest version of the reliability number tells you how much you can trust a single judge grading alone. But we never publish a single judge’s score; we publish the average of three judges, and that average is more trustworthy than any one judge on its own. By reporting both the reliability ratings of individual judges and the reliability of their average score, we produce a more robust assessment of the overall reliability of our AI-scoring process.
We also report something simpler alongside all this: an estimate of how many points a score could be off by. Agreement statistics have a hidden trap. ICC depends partly on the amount of variation among the cases being rated, so a diverse set of cases can produce a high ICC even if the actual grading is sloppy. Reporting a measure of absolute error alongside ICC makes the disagreement visible, allowing auditors to judge whether the magnitude of disagreement is acceptable for the intended use. Our goal is to produce accurate ratings of AI systems that are also fully transparent about the methods we use to produce those ratings.
Why this is rare, and why it matters
Most AI scores presented on leaderboards, in benchmark reports, or in glossy safety claims arrive as a bare number with nothing attached to tell you how much to trust it.
Although scholars increasingly recognize the importance of examining the reliability and bias of LLM-based judges, it is rarely reported alongside safety ratings. The measurement gets made; the trustworthiness of the measurement goes unmentioned. As AI grading AI becomes a normal part of how these systems get evaluated, evidence of reliability should accompany every score. Publishing the reliability is the standard we are arguing for, and the one we hold ourselves to.
Where agreement stops
Agreement is necessary but not sufficient. A high ICC tells you the judges are consistent. It does not tell you they are right.
Three models trained on overlapping data and tuned in similar ways can share the same blind spot, and a bias written into the prompts or the rubric will nudge all three judges toward the same result, producing tidy agreement on the wrong answer.
That is why we treat ICC as one diagnostic, not the final assessment of score quality. We interpret it alongside the size of the error, checks better suited to coarse scales, and, above all, expert human judgment, which remains the closest thing we have to the ground truth. Reliability is not the same as accuracy. Nevertheless, without adequate reliability, claims about accuracy or validity are difficult to interpret.
A score should be accompanied by evidence about agreement among the independent judges who produced it. That is not extra credit. It is what turns a number into evidence.
Read the methodology · Read the reliability paper
Further reading
- Shrout, P.E. & Fleiss, J.L. (1979). “Intraclass correlations: Uses in assessing rater reliability.” Psychological Bulletin, 86(2), 420–428. doi:10.1037/0033-2909.86.2.420
- Reuel-Lamparth et al. (2024). “BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices.” NeurIPS 2024. arXiv:2411.12990
- Verga et al. (2024). “Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models.” arXiv:2404.18796
- Bean, A.M. et al. (2025). “Measuring what Matters: Construct Validity in Large Language Model Benchmarks.” NeurIPS 2025. arXiv:2511.04703
- Gu, J., Jiang, X., Shi, Z., Tan, H. et al. (2024). “A Survey on LLM-as-a-Judge.” arXiv:2411.15594
- Shi, L., Ma, C., Liang, W., Diao, X., Ma, W. & Vosoughi, S. (2024). “Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge.” arXiv:2406.07791