GSS 2024: Preserving Uncertainty Improves Survey Fit
A target-separated GSS pilot compares four response conditions across 360 people and 17 questions, separating probability elicitation from profile value.
The GSS 2024 pilot reached 92.47% aggregate approximation with structured-profile probability answers. Its mean distribution error was 7.53 percentage points across 17 questions and 360 matched respondents. The useful finding was not simply a high final score: collecting probabilities substantially improved distribution fit over asking for one forced answer.
What was compared
The evaluation used real GSS attitudes and behavior responses as held-out targets. Four conditions distinguished the answer format from the context supplied to the model: generic forced choice, generic probability, demographic probability, and structured-profile probability. Candidates were sealed before the evaluation targets were opened.
The profile contained structured survey evidence, not the complete production Mind representation. The experiment ran in an isolated harness. Its result supports the tested response method, not identical accuracy for every production audience.
Results across the four conditions
Mean item absolute error
Reported results for this study. The table and discussion retain the study scope, baselines and uncertainty.
- Generic forced choice25.74
- Generic probability8.44
- Demographic probability7.67
- Structured-profile probability7.53
| Condition | Mean item absolute error |
|---|---|
| Generic forced choice | 25.74 pp |
| Generic probability | 8.44 pp |
| Demographic probability | 7.67 pp |
| Structured-profile probability | 7.53 pp |
Approximation is 100 minus mean absolute percentage-point error. The resulting 92.47% summarizes answer-distribution fit, not the percentage of individual respondents whose choices were predicted correctly. The number of response options and the aggregation method affect this measure.
The registered resampling estimate for generic probability versus forced choice was a 16.51-point reduction in error. This is the reported bootstrap estimate, not subtraction of the rounded descriptive table values. It isolates the response-format comparison; it must not be described as a profile or foundation-model superiority result.
What the profile comparison adds
Structured profiles had a small descriptive advantage over demographic probability. The registered profile lift was 0.17 points, with a 95% interval from -1.00 to +1.31. The interval crossed zero and the profile condition did not meet its promotion criterion. The study therefore has a mixed overall outcome even though probability elicitation supplied a useful positive finding.
The practical distinction is important. Better distribution fit from probability collection does not by itself demonstrate that adding biographical detail helps. A fair grounding comparison holds the answer format constant.
What teams can take from this result
When the target is a distribution of survey responses, preserving uncertainty can be more informative than reducing each synthetic respondent to a single categorical answer. Evaluation should retain both a matched generic probability baseline and a demographic baseline so that the incremental contribution of richer context remains visible.
This study does not establish a representative estimate for an arbitrary customer audience, exact individual prediction, or an error guarantee on new questionnaires. Withholding runtime targets also cannot prove that public data never appeared in model pretraining.
The evidence overview places GSS beside four other survey comparisons. The ANES follow-up tests later answers using earlier respondent evidence. The production Gen Z comparison evaluates a separate persistent-Mind cohort.


