How the AI Ethics Index Works
Turning responsible-AI principles into evidence.
Most AI oversight stops at what a company says it does. The AI Ethics Index looks at what tested products actually do. We translate broad ethical commitments—such as fairness, safety, and respect for privacy—into something you can measure: a specific, testable claim, checked against dated evidence.
The framework: one home for every question
A central goal of the AI Ethics Index is to score design affordances. An affordance is a property of a system’s design that enables, constrains, or makes specific user actions more likely. Rather than scoring marketing claims or intentions, we score what the product is built to allow.
Each claim and indicator is assigned to one location in a four-level hierarchy to reduce duplication and make aggregation traceable. Each issue has a defined location in the framework:
Nine dimensions organize the ethical, technical, and societal questions addressed by the Index, including Fairness, Privacy and Data Stewardship, and Human-AI Interaction.
Subcategories divide each dimension into more specific themes, such as User-Interface Transparency.
Assessable claims with defined evidence requirements, such as “the system supports user consent and control.”
The smallest scored units, each with explicit criteria and a specified evidence type, such as a right-to-exit or override control.
The nine dimensions
Model Design & Development
Examines whether objectives, assumptions, constraints, and intended uses are clearly documented and appropriately justified.
Fairness
Examines disparities in performance and impact, representational harms, and structural sources of bias using quantitative and contextual evidence.
Privacy & Data Stewardship
Examines provenance, consent pathways, retention, and re-identification risk.
Transparency
Examines clarity and completeness of documentation, disclosures, and known limitations.
Knowledge & Attribution
Examines factual accuracy, patterns of unsupported claims, source attribution, and the distinction between evidence and inference.
Human–AI Interaction
Examines usability, risk of overreliance, and differential effects across user groups.
Safety & Security
Examines adversarial resilience, jailbreak resistance, and safe-failure behavior under stress.
Societal Impact
Examines second-order effects on communities, institutions, labor, and democratic trust.
Governance & Accountability
Examines internal governance, incident response, versioning, and mechanisms for redress.
How we gather evidence
Different claims call for different kinds of evidence. Documentation can establish what a provider discloses; live tests can examine behavior; educational quality may require trained expert judgment. Each indicator specifies the evidence required and the method used to assess it.
Document analysis. For an indicator about whether a provider has documented a policy or commitment, the dated document is the primary evidence. A judge reads the relevant material, such as a model card or a children’s privacy notice, and answers a few focused questions: is the claim actually written down? Is it specific, complete, applicable, and internally consistent? When the document set is small enough, the complete materials are reviewed. Larger document sets are indexed for retrieval, and the passages used in scoring are retained with citations so that the evidence can be audited.
Live behavioral testing. When a claim is about behavior, the model’s own responses are the evidence, not what its documentation promises. For relevant behavioral indicators, we construct multiple conversation arcs that vary the use condition, escalation pattern, and conversation length. Each scenario is run through a standardized interaction protocol designed to approximate the intended use. Then a judge reads the whole exchange and scores it against a rubric. The indicator’s score is the average across its arcs.
Human judgment. Some assessments resist automation, including assessments of moral conduct, relational posture, and the quality of teaching. Here, trained human evaluators use a shared rubric, engage the live system the way a real user would, and rate the exchange. Three judges score each case, and how closely they agree helps us determine how much confidence to place in the score.
How we score
No published score is based on one judgment alone. AI-scored indicators are rated independently by three models from different model families, while human-scored indicators are rated by multiple trained evaluators. The judges review the same evidence, assign scores from 0.0 to 1.0, and provide written rationales. We then average their scores. A human reviewer also examines AI-panel results before publication.
Agreement is itself a signal. When the three scores land close together, the result is solid. When the scores diverge, that triggers a defined review procedure. We quantify agreement with a statistic called the intraclass correlation coefficient (ICC). Every reported score includes an inter-rater reliability statistic and an absolute measure of disagreement alongside the score. For every measurement, we produce both the average score and a measure of confidence in that score.
How indicators are built
Each indicator is developed through a documented process and includes a written specification, defined success criteria, and a stated evidence requirement. It is reviewed by internal and external specialists, pilot-tested, analyzed for reliability and validity, and revised before release. Indicators are versioned, and their use is logged. Changes and applications are documented so that the framework can be audited over time and continuously improved rather than quietly drifting.
Weights and lenses
Indicator scores are aggregated into claims, subcategories, and dimensions according to published rules, producing a structured profile that preserves both summary results and the underlying indicator scores for those who want to drill down. The output is an interpretable, multidimensional profile rather than a single rank.
The weights used in that aggregation are deliberately adjustable, and we treat them as lenses. A perspective that prioritizes community wellbeing will weigh some indicators differently than one centered on individual autonomy. A safety-first orientation weighs differently than an innovation-first one, and a long-term view differently than a short-term one. Lenses let the same evidence be read through different values without rewriting the framework, so a score always states the priorities it reflects.
Composite views
On top of the core dimensions, we publish composite views: outcome-focused versions assembled from indicators already in the framework, adding insight without increasing testing burden. They let a reader look at the theme they care about most—for example, Child and Youth Safety and Wellbeing, Educational Impact, or Human Flourishing—and see how a system performs there specifically.
Preventing gaming and drift over time
A standard is only as good as its resistance to gaming and drift. We use confidential evaluation sets, attested disclosures, and spot audits so a product cannot simply teach to the test. Confidential materials may be reviewed by authorized independent auditors under controlled access, while remaining unavailable to evaluated vendors. The framework itself is kept under version control, monitored for drift, reviewed when incidents occur, and audited on a regular schedule by an independent group of technical, ethical, and domain experts. The Index is intended for repeated evaluation across product versions rather than one-time certification.
A worked example: equity
Equity shows how the pieces come together. We ask whether an AI product delivers fair outcomes for everyone, not just for the users with the most resources, familiarity, or access.
In practice, we look for a few concrete things. Does the product perform as well for disadvantaged groups as it does for others? We check this directly, comparing error and acceptance rates across groups, and running counterfactual tests that change a single characteristic—such as language background or disability—in otherwise identical inputs to determine whether differences in the outputs are relevant and justified and whether the outputs offer equivalent quality and opportunity. Does it work for students who read in a language other than English, or who rely on assistive technology? Does the tool require devices, bandwidth, or prior technical fluency that some schools and families may not have? And because a tool can look fair on each measure on its own while failing groups defined by intersecting characteristics, we report results for those intersecting groups rather than folding them into an average.
What a score tells you
A score reflects how a product stood up to structured, independent testing on the claims we measured, and it comes with a measure of how much that result can be trusted. It covers what we tested, not every possible interaction, and it is built to inform rather than replace the judgment of the educators, families, and institutions who know their own context.