·Validation·Minds Team

PISA 2022: Synthetic Survey Fit Across Five Countries

A five-country PISA comparison evaluates 600 students and ten held-out questions, with weighted distribution errors and explicit country and language limits.

The PISA 2022 comparison reached 93.80% aggregate approximation using compact-profile probability answers. Weighted mean distribution error was 6.20 percentage points across ten held-out items from 600 students in five countries.

The repeatable positive finding concerned response format. Generic probabilities reduced weighted error by 14.86 points relative to generic forced choices, with positive effects across every evaluated country and question family.

The international comparison

The sample included 120 students from each of the United Kingdom, Germany, Japan, Mexico and the United States. Ten questions covered five question families. Conditions compared generic forced choice, generic probability, country-plus-demographic probability, and compact-profile probability.

The test used English master wording across all countries. It evaluated country context, not the performance of five localized-language questionnaires. Human evaluation outcomes were held out from runtime inputs.

Weighted distribution results

Weighted mean absolute error

Reported results for this study. The table and discussion retain the study scope, baselines and uncertainty.

  • Generic forced choice22.29
  • Generic probability7.42
  • Country and demographic probability6.59
  • Compact-profile probability6.20
0Percentage points · lower is better30
ConditionWeighted mean absolute error
Generic forced choice22.29 pp
Generic probability7.42 pp
Country and demographic probability6.59 pp
Compact-profile probability6.20 pp

Approximation is 100 minus weighted mean absolute percentage-point error. The 93.80% result is a transformation of the 6.20-point error, not a claim that 93.80% of students were individually predicted correctly.

The registered weighted probability lift was 14.8631 points, rounded to 14.86, with a 95% interval from 8.78 to 18.96. This contrast concerns probability collection versus forced answers within the generic model condition. It must not be described as Minds profiles outperforming another foundation model by that amount.

Country coverage is not universal profile value

The compact profile had a small descriptive advantage over country and demographics. The registered lift was 0.40 points, with an interval from -0.53 to +1.47, leaving the incremental aggregate effect inconclusive.

The profile condition failed its country and question-family promotion gate and worsened individual Brier score and exact accuracy relative to demographics. The study therefore has a mixed outcome despite its positive probability-elicitation finding.

That boundary matters when interpreting an international result. A favorable weighted average does not imply that a richer profile improves every country's result or every question family.

Where this evidence applies

The experiment supports keeping uncertainty in synthetic answers when the target is a distribution. Its country-level consistency strengthens the response-format finding beyond one national dataset.

It does not produce official OECD population estimates, validate every educational topic, or demonstrate multilingual prompting accuracy. The isolated-harness result also does not guarantee identical performance on every current product path. Withholding runtime targets cannot exclude exposure to public data during model pretraining.

For a customer decision, the relevant question remains whether the instrument, audience, language and scoring method match the proposed use. These differences should be checked before treating a benchmark score as an expected result.

The CES questionnaire comparison covers several response forms. The Gen Z production study tests another audience through the product. The evidence overview keeps international, product and qualitative evidence distinct.

Reference data

PISA 2022