When Synthetic Audiences Outperform Generic Foundation Models
A held-out interview comparison tests when participant profiles and memory improve synthetic responses over generic GPT, Claude and Gemini prompts.
A foundation model can produce a plausible answer to an interview question. Reproducing what a particular person would say is a more specific task. Their past experiences, values and reasons may matter more than the model's ability to produce a fluent general response.
We tested that distinction using 66 held-out interview answers from 22 people. A synthetic respondent with a compact participant profile and retrieved earlier evidence outperformed generic GPT-5.4 and Claude Sonnet 5 responses on a blinded participant-fit composite. The primary advantage was 25.49 points over GPT and 30.05 points over Claude. A second judge version reproduced the positive comparisons.
The comparison supports a specific use of synthetic audiences: bringing relevant participant evidence into questions that depend on that evidence. The generic models received no participant history. This is therefore a comparison of a grounded workflow with generic prompting, rather than a ranking of model capability under identical context.
The test: earlier evidence, later answers
The benchmark used two public interview corpora covering laser-dentistry practitioners and refugee food insecurity. Earlier interview evidence became compact profiles and retrievable memories. Three later answers per person remained held out for evaluation, giving 66 target answers from 22 people.
Each target question had four anonymous candidate responses: profile plus retrieval, a generic response from the same underlying Gemini-family model, generic GPT-5.4, and generic Claude Sonnet 5. Candidate answers were sealed before scoring and judges evaluated them without condition labels.
The rubric assessed alignment with the real answer, viewpoint and values, reasons, specificity, voice and length fit, alongside genericness and unsupported personal facts. These are dimensions of participant fit. They are not measures of ad conversion, general intelligence or the percentage of humans correctly simulated.
The measured advantage
Participant-fit composite
22 people, 66 held-out answers. Primary automated judge. Generic models did not receive participant history; this is not an equal-context model ranking.
- Profile + retrieval67.66
- Generic GPT-5.442.17
- Generic Claude Sonnet 537.61
- Generic Gemini35.24
| Comparison with profile plus retrieval | Primary composite lift | Participant-clustered 95% bootstrap interval | Second-judge lift |
|---|---|---|---|
| Same-model generic Gemini response | +32.42 points | +23.93 to +40.52 | +32.83 points |
| Generic GPT-5.4 | +25.49 points | +15.99 to +34.56 | +26.13 points |
| Generic Claude Sonnet 5 | +30.05 points | +21.86 to +38.12 | +30.68 points |
Under the primary judge, the grounded condition scored 67.66 on the composite, compared with 42.17 for generic GPT, 37.61 for generic Claude and 35.24 for the same-model generic response. Both corpora showed positive average advantages over the named GPT and Claude conditions under both judge versions.
The two judge versions agreed on the winner for 81.8% of items. Both judges were Gemini-family models, so this checks version sensitivity within that family, not agreement between independent human raters or unrelated judge families.
Why grounding can help
The most direct interpretation is access to relevant evidence. A generic model has to fill in missing personal context using general patterns. A grounded respondent can instead draw on recorded experiences and preferences.
For a question about a professional decision, earlier evidence may identify the person's constraints, prior experience and reason for choosing one option over another. The answer can then reflect those specifics instead of giving broadly reasonable advice. In the benchmark, the grounded condition received higher specificity and reasons scores and was judged less generic.
The same-model generic comparison is useful here. Keeping the underlying model family fixed while changing the profile-and-retrieval condition shows an advantage for the added grounding workflow on this task. It does not isolate which part of that workflow caused how much of the gain. Profiles, retrieval, extra context and associated prompting changed together.
The results are consistent with the value of relevant memory. They do not establish that adding more context always helps, or that every stored fact deserves to appear in every answer.
When the advantage is most plausible
This evidence is most directly relevant when a research question depends on personal history, a previously expressed preference, a concrete experience or the reasons behind a decision. Follow-up interviews and exploration of why audience members interpret the same proposition differently fit that pattern.
A useful practical test is whether knowing this particular person's earlier answers could materially change the response. If so, grounded evidence may be valuable. If a question only requires general factual knowledge, this experiment does not establish an advantage for a synthetic audience over a capable generic model.
Even within the interview benchmark, the advantage was not universal. In the exploratory reflective-question slice, GPT scored 10.75 points above the grounded condition. That slice contained only five items from three people, so it is too small to establish a reliable task ranking. It nevertheless prevents the average win from being interpreted as a win on every kind of question.
Surveys require a separate explanation
The Against Reality survey overview found that probability-based response collection consistently improved aggregate distributions over forced-choice answers. Detailed profiles did not consistently improve those distributions beyond matched demographic probability prompts.
That matters for a fair foundation-model comparison. Comparing a probability-based audience with a forced-choice generic model mixes the effect of grounding with the effect of response format. Both should receive the same answer contract when the purpose is to measure incremental grounding value.
The Gen Z food-survey test provides a narrower quantitative example. Its later production run had 6.01-point error versus 6.68 for a frozen generic probability baseline. The 0.67-point difference is descriptive because the generic baseline was historical. It cannot carry a broad claim that synthetic audiences beat foundation models on surveys.
What this comparison leaves open
The study is small: two corpora, 22 people and 66 answers. Some interviews were translated into English, affecting what voice and style mean. Model versions and sampling controls differed, and the named generic models had no participant evidence. An equal-context comparison would need to give each model the same permitted source material, hold the response contract fixed and report the resulting cost as well as quality.
The scoring also reused the interview evidence set from an earlier confirmation. That earlier comparison and this four-condition evaluation are not two independent samples. Future replication should include new participants, different subject areas, independent human ratings and stronger controls for which information each system sees.
The claim this study supports is useful without being universal: on these held-out interview questions, a participant-grounded synthetic audience produced answers that better matched the real people than generic responses from the tested model versions. Relevant evidence was central to the advantage.
For the workflow behind audience grounding, see how Minds panels are built. The validation checklist describes how to connect a synthetic result to the evidence required for a decision.


