---
title: "Minds vs GPT-5.4 and Claude Sonnet 5 on Real… | Minds"
canonical_url: "https://getminds.ai/research/minds-vs-foundation-models-qualitative-2026"
last_updated: "2026-08-13T13:07:34.539Z"
meta:
  description: "A blinded qualitative benchmark comparing evidence-grounded Minds with generic GPT-5.4 and Claude Sonnet 5 on held-out human interview answers."
  "og:description": "A blinded qualitative benchmark comparing evidence-grounded Minds with generic GPT-5.4 and Claude Sonnet 5 on held-out human interview answers."
  "og:title": "Minds vs GPT-5.4 and Claude Sonnet 5 on Real… | Minds"
  "twitter:description": "A blinded qualitative benchmark comparing evidence-grounded Minds with generic GPT-5.4 and Claude Sonnet 5 on held-out human interview answers."
  "twitter:title": "Minds vs GPT-5.4 and Claude Sonnet 5 on Real… | Minds"
---

Minds

August 11, 2026·Validation·Minds Team

# **Minds vs GPT-5.4 and Claude Sonnet 5 on Real Interview Answers**

Across 66 later interview answers from 22 real people, Minds beat generic GPT-5.4 by 25.49 points and Claude Sonnet 5 by 30.05 points, replicated by a second judge.

[Run qualitative research](https://getminds.ai/?register=true)

What makes a synthetic respondent useful is not whether it can produce a plausible quote. It is whether it can continue the perspective of a real person when asked a new question.

Minds Research Lab tested that directly. We used 66 held-out later answers from 22 real interviewees across two independent public qualitative datasets. Evidence-grounded Minds were compared with generic GPT-5.4, Claude Sonnet 5, and a same-model generic response.

Minds beat GPT-5.4 by **25.49 composite points** and Claude Sonnet 5 by **30.05 points**. A second blinded evaluator reproduced both results.

**25.49** pts

Lift vs GPT-5.4

**30.05** pts

Lift vs Claude Sonnet 5

**66**

Held-out human answers

## The benchmark

The study used two independently published, openly licensed interview corpora:

- [De-identified interviews about refugee food insecurity](https://doi.org/10.6084/m9.figshare.29318549.v1), published by University of Utah researchers.
- [International laser-dentistry interview transcripts](https://doi.org/10.6084/m9.figshare.28456946.v1), published by Khon Kaen University researchers.

For every participant, earlier material could inform the synthetic respondent while later answers were held back for evaluation. The models received the same interview question and the same answer-form instruction. Generic foundation models received no participant memory or profile.

All candidates were generated before the held-out answers were used for scoring. Candidate labels were randomized so the evaluator could not see which model produced which answer.

## Results

| Candidate | Composite score | Winner rate | Difference versus Minds |
| --- | ---: | ---: | ---: |
| **Minds** | **67.66** | **72.7%** | - |
| GPT-5.4 | 42.17 | 16.7% | **Minds +25.49** |
| Claude Sonnet 5 | 37.61 | 0.0% | **Minds +30.05** |
| Generic same-model response | 35.24 | 3.0% | **Minds +32.42** |

The participant-clustered 95% interval was **+15.99 to +34.56** versus GPT-5.4 and **+21.86 to +38.12** versus Claude Sonnet 5. Both interview corpora were positive independently.

A second evaluator version reproduced lifts of **+26.13** and **+30.68** points. The two judges agreed on the winning candidate for 81.8% of individual answers.

## Where Minds gained

The rubric measured more than surface fluency. It compared each candidate with the person's real later answer across:

- Semantic alignment
- Viewpoint and values
- Themes and reasons
- Specificity
- Voice and style
- Length fit
- Genericness
- Unsupported personal facts

Minds' largest advantages appeared in semantic alignment, viewpoint, reasons, specificity, and recognizable voice. In other words, Minds did not merely sound more polished. It was more likely to preserve what that person cared about, how they explained it, and which details they naturally used.

The advantage was not a length artifact. Candidate answers averaged roughly 52 to 60 words across all conditions.

## Why this differs from asking ChatGPT to act like a persona

A generic model is exceptionally capable, but the question alone does not contain the person's history. It must produce an answer that is plausible for someone in the abstract.

A Mind is built to maintain a persistent research participant. It can carry forward approved knowledge, prior research context, and a differentiated point of view across new questions. That allows a smaller or lower-cost model to outperform a much larger generic model when the task depends on knowing the person rather than knowing the world.

This is the central Minds thesis in measurable form: **personality depth and relevant memory can matter more than raw foundation-model scale.**

## Additional qualitative validation

The same broader research program also tested:

- [PRISM](https://arxiv.org/abs/2404.16019), using 299 untouched participants and 869 eligible human responses across values, controversial topics, and unguided prompts.
- The [NASA Oral History collection](https://www.nasa.gov/history/history-publications-and-resources/oral-histories/nasa-list/), using later answers to test continuity across long-form professional interviews.
- A separate 21-person language-cohort holdout from the refugee interview corpus.

Across PRISM, complete Minds profiles scored approximately **10 to 12 points above** generic and demographic prompting. In the separate language-cohort holdout, the Mind condition improved by **20.77 points** over generic responses overall.

## What this means for customer research

Generic foundation models are useful for brainstorming plausible reactions. Minds is designed for a different standard: preserving differentiated participants across a study.

That makes Minds especially relevant for exploratory interviews, concept reactions, message testing, objection discovery, buying-committee research, persona-based follow-ups, and turning an existing research archive into an interactive audience.

The named-model result covers the tested interview sources and model versions. Minds Research Lab continues to extend the benchmark across domains, languages, and question forms while keeping the core comparison anchored in real human answers.

For the quantitative evidence, read [Real Survey Benchmarks](https://getminds.ai/research/survey-approximation-benchmarks-2026). For the complete program, read [Minds Research Benchmark 2026](https://getminds.ai/research/real-world-validation-benchmarks-2026).

## Public benchmark sources

- [Utah refugee interview transcripts](https://doi.org/10.6084/m9.figshare.29318549.v1)
- [International laser-dentistry transcripts](https://doi.org/10.6084/m9.figshare.28456946.v1)
- [PRISM Alignment Dataset](https://arxiv.org/abs/2404.16019)
- [NASA Oral Histories](https://www.nasa.gov/history/history-publications-and-resources/oral-histories/nasa-list/)

## Suggested citation

Minds Research Lab. “Minds vs GPT-5.4 and Claude Sonnet 5 on Real Interview Answers.” August 11, 2026. https://getminds.ai/research/minds-vs-foundation-models-qualitative-2026

## **Frequently asked questions**

### **Why did Minds beat larger foundation models?**

The generic models had to infer a plausible answer from the interview question alone. Minds responded with persistent, evidence-grounded participant context, allowing it to preserve the person's viewpoint, reasons, specificity, and voice.

### **Was the result evaluated blindly?**

Yes. Candidate answers were anonymously ordered, evaluated against held-out human answers, and scored with the same rubric. A second blinded evaluator reproduced the ranking.

### **Was the Minds advantage caused by longer answers?**

No. The compared answers averaged approximately 52 to 60 words. Minds gained primarily in semantic alignment, viewpoint, reasons, specificity, and recognizable voice.