·Validation·Minds Team

We Tested Synthetic Audiences Against Reality

What five public survey datasets reveal about synthetic audience accuracy, including a 301-Mind Gen Z production test, and where profile grounding helps.

Synthetic research needs an external reference: answers from real people. Our original research suite compared synthetic answers with four public survey datasets and two qualitative benchmarks. A later Food and You production test added a practical check on the same question: how closely can an audience of persistent Minds reproduce a real survey's answer distributions?

The results point to two different sources of value. Preserving uncertainty improves survey-distribution estimates. Participant profiles and relevant memories provide additional value when an answer depends on someone's experiences, reasons and values. Their benefits vary by task.

Selected survey replications

Mean distribution error

Different questionnaires and cohorts; PISA is weighted. Lower error is better. Do not pool these results into one accuracy score.

  • GSS 20247.53
  • ANES 20248.33
  • CES 20244.72
  • PISA 20226.20
  • Food and You 2, Wave 106.01
0Percentage points · lower is better10
Replicated studyPublic source datasetOutcome
GSS 2024 survey replicationGSS 202492.47% approximation · 7.53 pp mean error
ANES 2024 longitudinal replicationANES 2024 Time Series Study91.67% approximation · 8.33 pp mean error
CES 2024 pre/post replication2024 Canadian Election Study95.28% approximation · 4.72 pp mean error
PISA 2022 international replicationPISA 2022 Database93.80% approximation · 6.20 pp weighted mean error
Food and You 2 production replicationFood and You 2, Wave 1093.99% approximation · 6.01 pp mean error · Pearson 0.804 · Spearman 0.834

Aggregate approximation means 100 minus mean absolute percentage-point error. It measures the gap between answer distributions. It does not mean that the same percentage of individuals was predicted correctly. The studies use different instruments and aggregation details, including a weighted PISA result; the rows should not be averaged into one accuracy score.

The first four rows belong to the original isolated survey experiments. The fifth is a later product run on an existing locked cohort. Its inclusion extends the evidence overview without turning a repeat on that cohort into a new independent population study.

Gen Z: a real questionnaire through the product

The UK Food Standards Agency's Food and You 2 survey asked how often young people eat out or buy takeaway food for breakfast, lunch and dinner. The production comparison asked the same three questions to 301 persistent Minds aged 16 to 24 and compared the resulting distributions with the published youth results. The human question base was 257 for each evaluated item.

All 903 planned responses completed. Across 21 answer-option cells, mean absolute error was 6.01 percentage points and mean distribution overlap was 78.98%. These describe different properties: the approximation transformation summarizes average cell error, while overlap shows how much probability mass the distributions share.

The August 18 run improved on the August 9 result using the same cohort and instrument: error fell from 7.33 to 6.01 points, an 18.0% relative reduction. Every question improved. The comparison was a product regression test with a historical control, rather than a concurrent randomized comparison.

Baseline context matters. A frozen generic probability prompt recorded 6.68-point error and an equal-probability null recorded 7.00 points on the same cells. The production result was better descriptively, but those baselines were not rerun contemporaneously. A score in the nineties on the approximation scale therefore should not be read as a large advantage by itself.

The full Gen Z food-survey article describes the audience, questionnaire, provenance and per-question results. The Food Standards Agency supplied the public reference dataset; it did not sponsor, endorse or review this Minds study.

Why uncertainty improved survey estimates

In the four original survey experiments, probability answers reduced distribution error by 12.47 to 19.72 percentage points relative to forcing one answer. The improvement appeared across all four datasets.

A probability answer retains uncertainty that a categorical choice discards. When aggregated across an audience, these probabilities can recover a population distribution more closely than one forced choice from each synthetic respondent.

That result is about response collection and aggregation. The same experiments did not show that detailed profiles consistently beat demographic probability prompts. The comparison baseline and response format matter as much as the final score.

Other selected validations

These studies use different outcomes, so an approximation percentage would be misleading. They complement the survey evidence rather than enlarge its accuracy range.

Study and detailed resultsEvaluated sampleReported resultApproximation
PRISM Alignment299 participants, 869 items32.6% participant-fit winner rate; 25.7% demographic and 20.5% genericNot applicable
Held-out interviews22 people, 66 answers, two corpora+24.37 composite points over generic, participant-clustered estimateNot applicable
Named foundation modelsSame 22 people and 66 answers, not a new cohort+25.49 points over GPT-5.4; +30.05 over Claude Sonnet 5Not applicable
Spark creative diversitySeven creative tasksMean diversity 7.90/10 versus 3.14/10 uniform baselineNot applicable

Spark evaluates diversity of ideas, not agreement with a representative survey or prediction of real advertising performance. The interview comparisons use automated judges, not a percentage of correctly simulated people.

Where participant context helped

The original qualitative tests assessed later human answers that had been withheld from generation. In the PRISM Alignment comparison, full Minds achieved a 32.6% blinded winner rate, compared with 25.7% for demographic prompts and 20.5% for generic prompts. This PRISM dataset is an external benchmark, distinct from Minds' own PRISM approach.

The independent interview confirmation used two public corpora, 22 people and 66 later answers. Its report found a positive profile-plus-retrieval advantage over generic prompting across both corpora and under a second judge version. A named-model comparison explores that advantage with generic GPT and Claude responses. Those baselines did not receive participant history.

These experiments concern participant fit, reasons, specificity and voice. Their judge scores should not be combined with survey-distribution percentages. Both depend on what evidence the synthetic respondent receives, and the qualitative assessments still require independent human-rater confirmation.

How far the evidence travels

The original suite comprised 1,881 participants across its survey and qualitative benchmarks. The additional 301-Mind food run has its own cohort and human reference denominator; those should remain separate.

Targets were withheld from runtime inputs in the reported tests. Public-source data may still have appeared in model pretraining, so runtime separation cannot establish absence of all prior exposure. The original survey results also came from isolated harnesses; they do not establish identical behavior across every product version. PISA used English master wording across its five countries.

This overview presents the approved evidence covered here, not a census of every experiment Minds has run. Other exploratory and failed tests require their own context. The practical use of these findings is to choose an appropriate audience, response format and validation method for a specific question.

Read how Minds research panels are built for the workflow, and use the validation checklist when deciding what a synthetic result can support.