---
title: "When Synthetic Audiences Outperform Generic… | Minds"
canonical_url: "https://getminds.ai/research/synthetic-audiences-vs-foundation-models"
last_updated: "2026-09-08T17:04:39.491Z"
meta:
  description: "A held-out interview comparison tests when participant profiles and memory improve synthetic responses over generic GPT, Claude and Gemini prompts."
  "og:description": "A held-out interview comparison tests when participant profiles and memory improve synthetic responses over generic GPT, Claude and Gemini prompts."
  "og:title": "When Synthetic Audiences Outperform Generic… | Minds"
  "twitter:description": "A held-out interview comparison tests when participant profiles and memory improve synthetic responses over generic GPT, Claude and Gemini prompts."
  "twitter:title": "When Synthetic Audiences Outperform Generic… | Minds"
---

Minds

September 5, 2026·Research·Minds Team # **When Synthetic Audiences Outperform Generic Foundation Models** A held-out interview comparison tests when participant profiles and memory improve synthetic responses over generic GPT, Claude and Gemini prompts. A foundation model can produce a plausible answer to an interview question. Reproducing what a particular person would say is a more specific task. Their past experiences, values and reasons may matter more than the model's ability to produce a fluent general response. We tested that distinction using 66 held-out interview answers from 22 people. A synthetic respondent with a compact participant profile and retrieved earlier evidence outperformed generic GPT-5.4 and Claude Sonnet 5 responses on a blinded participant-fit composite. The primary advantage was 25.49 points over GPT and 30.05 points over Claude. A second judge version reproduced the positive comparisons. The comparison supports a specific use of synthetic audiences: bringing relevant participant evidence into questions that depend on that evidence. The generic models received no participant history. This is therefore a comparison of a grounded workflow with generic prompting, rather than a ranking of model capability under identical context. ## The test: earlier evidence, later answers The benchmark used two public interview corpora covering laser-dentistry practitioners and refugee food insecurity. Earlier interview evidence became compact profiles and retrievable memories. Three later answers per person remained held out for evaluation, giving 66 target answers from 22 people. Each target question had four anonymous candidate responses: profile plus retrieval, a generic response from the same underlying Gemini-family model, generic GPT-5.4, and generic Claude Sonnet 5. Candidate answers were sealed before scoring and judges evaluated them without condition labels. The rubric assessed alignment with the real answer, viewpoint and values, reasons, specificity, voice and length fit, alongside genericness and unsupported personal facts. These are dimensions of participant fit. They are not measures of ad conversion, general intelligence or the percentage of humans correctly simulated. ## The measured advantage_### **Participant-fit composite** 22 people, 66 held-out answers. Primary automated judge. Generic models did not receive participant history; this is not an equal-context model ranking._- Profile + retrieval**67.66**