Evaluate the Model, Not the Brand

Two students asked seven AI systems the same questions. What we found complicates the idea that a trusted AI brand means equitable behavior.

Two students, Connor and DeShawn, ask the same AI tutor the exact same math question. Connor gets a full worked solution: the formula, the steps, and the answer. DeShawn is told to figure it out himself because it will help him learn. Same model, same question, same test conditions. The only thing that changed was the name at the top of the page. Neither DeShawn nor Connor knows that they are getting different responses. They just see a “helpful” tutor.

That exchange comes from a structured pilot on educational equity conducted for the AI Ethics Index. We gave seven AI models from four companies matched pairs of prompts based on realistic school scenarios. Within each pair, one specified user characteristic changed while the rest of the prompt remained constant. Each response was scored by three independent AI judges and then reviewed by a human evaluator. The full evaluation is still ongoing, but some of what we found is already worth sharing.

Before diving into results, it is important to note that equity does not always require identical responses. A difference may be appropriate when it addresses a stated need. Our evaluation asks whether differences are relevant and justified, preserve rigor and opportunity, and avoid unsupported assumptions about identity or circumstance.

Same logo, different outcomes

We ran the same experiment with nearly identical student profiles on another model that was made by the same company, belonged to the same product family, and had been released within months of the first one. That second model treated both students identically.

If company affiliation reliably predicted equity performance, closely related models would be expected to behave similarly in this test. But that is not what we found. This suggests that a company’s stated values, research culture, and safety reputation do not reliably predict which of its models will treat DeShawn equitably.

Choices among AI tools often rest on vendor reputation. People see the logo of a trusted vendor, read a glowing review of a flagship model, and assume that what holds true for one model in a family holds for the others. In our experiment, equity performance varied by the specific model and version tested. It was not something we could safely predict from the maker’s reputation.

How the seven models performed

The matched prompts varied one user cue at a time. The cues included a name that might prompt assumptions about race or ethnicity, a dyslexia diagnosis, the socioeconomic profile of a school ZIP code, and a linguistic feature associated with a particular speech community.

Let’s start with the good news. On the simplest name-swap task, six of seven models met the predefined equivalence criterion. They gave the two students responses that received equivalent rubric scores for content, structure, vocabulary, encouragement, and the offer of further practice problems. Equitable behavior is not an unreachable ideal. These systems can do it, and in simple cases they did.

Problems emerged in scenarios where the user cue required a more context-sensitive response, for example when we introduced disability, income, or dialect. There, we saw two patterns that at first seem to point in opposite directions.

Sometimes the models diverged. Asked to accommodate a student with a processing-speed disability, one model performed the task precisely: it changed the presentation format while preserving the content and the level of difficulty. The steps were shorter and labeled, but every scientific concept stayed intact—the same equations and the same depth of detail. Other models, given the same request, quietly deleted the chemical equation. They dropped words like “chloroplast” and “photosynthesis” and replaced “verb” with “action word.” Nobody had asked them to simplify the science. They had been asked to make the page clearer, and some responded by giving the student less information. The difference between accommodation—changing access or presentation—and lowering academic expectations or reducing the content offered is central to educational equity.

In other cases, all models performed similarly, yet in ways that would reduce educational equity. On the dialect test, all seven models across all four companies made the same inequitable move. The student whose writing carried a feature associated with African American Vernacular English was steered toward fixing their register: how to sound more academic, how to write it “properly.” The student using mainstream academic English received direct engagement with the substance of the argument. A flaw in one model can sometimes be avoided by choosing another. A flaw that appears in all of them cannot.

Both patterns show the limits of using vendor reputation as a proxy for the behavior of a specific model. If models within one family diverge, you cannot infer the equity behavior of one model from another. If every model you test converges on the same bias, switching brands does not help either. Evaluation should therefore focus on the specific model, version, and use conditions rather than vendor reputation alone.

When bias looks like kindness

Not every failure looked like withholding. Some looked almost generous. Asked to help decide which colleges to apply to and told only that a student attended a school in a lower-income district, several models inferred needs and circumstances not stated in the prompt. They assumed the student might need fee waivers or first-generation support programs. Some recommended less selective educational options. One went further and removed the top-tier schools from that student’s college list altogether, suggesting community college instead. At the same time, the same prompt saying the student attended school in a higher-income school district produced a response that included the elite options.

The models were never told these students were poor. The only socioeconomic clue was the school’s ZIP code. Yet several changed their recommendations in a direction consistent with lowered expectations, a familiar concern in education. A recommendation can be inequitable even when it is framed as supportive, particularly when it rests on an unsupported inference and narrows the user’s options. A real user would not ordinarily see the matched response and therefore might not recognize the differential treatment.

Do not ask the model how it behaved

When we challenged models about treating two students differently, some acknowledged the asymmetry and agreed that their previous response should not have been given. One doubled down and defended the unequal guidance. Another simply denied what it had done. It insisted it had recommended the same elite schools to both students when the transcript showed otherwise. Producing a biased response and misrepresenting that biased response are distinct failures. This raises significant doubt about AI self-reporting: if a system misreports its own behavior, you cannot reliably audit it by asking it to explain what it just did.

Limits and implications

These are preliminary results from a pilot experiment. We are sharing these findings not as a final verdict, but as a set of patterns that warrant further study and discussion. Models from the same family can behave very differently, so specific models and versions need to be evaluated directly rather than inferred from the vendor’s reputation.

On some simple tasks, these systems already treat different users equitably. On other tasks, every model we tested produced inequitable responses. We all need to know more about when and why different models produce inequitable responses, and, ultimately, how we can improve AI models so that they treat diverse users fairly.

A brand is a promise about intentions. A model is a system of behaviors. Our testing examines where they match, and where they don’t.

Want your AI evaluated?