CES 2024: Survey Fit Across Multiple Question Forms
A CES comparison evaluates 300 matched respondents on 27 binary, ordinal and multiselect questions, separating aggregate fit from individual prediction.
The CES 2024 test reached 95.28% aggregate approximation with profile-probability answers, equivalent to 4.72 percentage points of mean distribution error. Across the response-format comparison, probability elicitation reduced error by a registered bootstrap estimate of 19.72 points relative to forced answers.
The study extends the survey evidence beyond a single question form. It also shows why a final approximation score must be accompanied by matched baselines.
The evaluated questionnaire
The comparison used 300 matched CES respondents and 27 held-out questions: 15 ordinal, 10 binary and two multiselect batteries. Earlier respondent information provided context; the evaluation answers remained held out from generation.
The four conditions were generic forced choice, generic probabilities or marginals, demographic probability, and full-profile probability. Multiselect options require marginal probabilities because a respondent may choose more than one option. Their interpretation differs from a single-choice probability vector.
The measured distributions
Mean item absolute error
Reported results for this study. The table and discussion retain the study scope, baselines and uncertainty.
- Generic forced choice26.927
- Generic probability or marginals7.036
- Demographic probability4.563
- Full-profile probability4.720
| Condition | Mean item absolute error |
|---|---|
| Generic forced choice | 26.927 pp |
| Generic probability or marginals | 7.036 pp |
| Demographic probability | 4.563 pp |
| Full-profile probability | 4.720 pp |
Approximation is 100 minus mean absolute percentage-point error. The 95.28% figure therefore describes average distribution fit on these evaluated questions. It is not a percentage of accurately simulated people.
The registered bootstrap probability lift was 19.719 points, rounded to 19.72. The reported point contrast was 19.892 points, with a 95% interval from 16.257 to 23.401. The bootstrap summary and the displayed condition means use different summaries and rounding; they should not be silently substituted for one another.
What richer profiles did and did not improve
Full profiles improved on generic probability prompting but did not beat matched demographic probability on the primary aggregate metric. The profile-versus-demographic lift was -0.157 points and its interval crossed zero.
Individual accuracy improved with profile context while aggregate error was slightly worse than with demographics. The result is therefore mixed, with a strong response-format finding but no confirmed aggregate advantage for the full-profile condition over the demographic condition.
Only two multiselect batteries were available. This is evidence from a mixed questionnaire, not universal validation of multiselect calibration across domains and instruments.
Applying the evidence
For distribution estimates, response collection and aggregation should be treated as explicit parts of the research design. A forced categorical baseline cannot isolate the value of profiles when the richer audience is allowed to return probabilities.
A comparison should also state whether its target is item-level population fit or individual prediction. More context is not automatically better for both. The tested sample, model configuration and questionnaire define the claim boundary.
This isolated-harness test does not guarantee current product behavior or representative estimates for new customer audiences. Public-source pretraining exposure remains possible even with runtime target separation.
The ANES longitudinal study examines the same aggregate-versus-individual distinction. The PISA comparison adds a five-country setting. The evidence overview connects these results without pooling them into one score.


