ANES 2024: Comparing Synthetic and Later Human Answers
An ANES pre/post-election comparison tests 300 matched people on 28 later answers, with separate results for response format and participant grounding.
The ANES 2024 comparison reached 91.67% aggregate approximation, or 8.33 percentage points of mean distribution error, using retrieved-profile probability answers. Its clearest measured improvement came from probability elicitation: generic probability answers reduced error by 12.47 points relative to generic forced choices.
Unlike a comparison that reconstructs answers already supplied to a model, this design used earlier pre-election respondent evidence and evaluated later held-out post-election answers.
The longitudinal test
The sample contained 300 matched respondents and 28 later items. Human targets were withheld from runtime generation and candidates were sealed before evaluation. Conditions separated a generic forced answer, generic probabilities, demographic probabilities, and probabilities informed by retrieved profile evidence.
Earlier evidence can provide context about a respondent, but it is not the same as knowing their later answer. The temporal separation makes this a useful test of whether prior information travels to a subsequent questionnaire.
Distribution results
Mean item absolute error
Reported results for this study. The table and discussion retain the study scope, baselines and uncertainty.
- Generic forced choice22.782
- Generic probability10.312
- Demographic probability8.263
- Retrieved-profile probability8.330
| Condition | Mean item absolute error |
|---|---|
| Generic forced choice | 22.782 pp |
| Generic probability | 10.312 pp |
| Demographic probability | 8.263 pp |
| Retrieved-profile probability | 8.330 pp |
The final approximation transformation is 100 minus 8.33, giving 91.67%. It describes average answer-cell distribution error, not individual forecasting accuracy and not a forecast of election outcomes.
For the generic probability versus generic forced-choice comparison, the registered improvement was 12.47 percentage points. No evaluated item regressed by more than two points on that contrast. The response-format result therefore extended beyond an isolated item.
Grounding and aggregation are different questions
Retrieved profiles did not improve aggregate mean absolute error over demographic probability prompts. The registered difference was +0.046 points of error, with an interval crossing zero. The descriptive table likewise shows demographic probability slightly ahead.
Profiles modestly improved individual Brier score and exact accuracy while slightly worsening population calibration. These outcomes can coexist: predicting a person's answer and reconstructing an audience's distribution are different targets. Selecting only the favorable metric would misstate the study.
The experiment's overall result is mixed. It supports probability elicitation while leaving the primary incremental aggregate benefit of profiles unconfirmed.
Implications and limits
Teams should select their evaluation target before comparing synthetic audiences. An individual-response task calls for individual metrics; an audience-distribution task needs aggregate calibration. A good score on one cannot substitute for the other.
This was an isolated harness using structured pre-election evidence, not the complete trained production Mind path. The evaluated sample and later questions do not establish general validity for every political audience or survey instrument. Runtime holdout does not rule out prior exposure to public data in model training.
The GSS pilot provides another response-format comparison. The CES study extends the instrument to several question forms. The evidence overview keeps their scores and scopes separate.


