·Validation·Minds Team

Held-Out Interviews: When Participant Memory Helps

Two public interview corpora test whether compact profiles and retrieved earlier evidence improve fidelity to 66 later answers from 22 real participants.

A compact profile with relevant retrieved evidence improved participant-fit scores over generic prompting in a confirmation using 22 people and 66 held-out interview answers. The registered participant-clustered lift was 24.37 composite points, with positive average advantages in both evaluated corpora.

The question was specific: can a grounded synthetic respondent better reflect what a particular person later said, rather than merely produce a plausible answer to an interview question?

Earlier evidence, later answers

The study used two public interview corpora concerning laser-dentistry practitioners and refugee food insecurity. Earlier interview material supplied compact participant profiles and retrievable evidence. Three later answers per person remained held out for scoring.

The same-model conditions distinguished generic prompting, role context, profile context, and profile plus retrieval. These comparisons help separate the value of a broad role description from the value of participant-specific evidence.

There were 66 target answers, but only 22 people. Multiple answers from one person are related observations; uncertainty therefore needs to respect the participant as a cluster.

The observed fidelity advantage

Composite result

Reported results for this study. The table and discussion retain the study scope, baselines and uncertainty.

  • Compact profile plus retrieval, mean60.33
  • Generic prompting, mean35.98
0Score / 100100
Condition or contrastComposite result
Compact profile plus retrieval, mean60.33
Generic prompting, mean35.98
Registered lift over generic prompting+24.37 points
Registered lift over role context+18.58 points
Registered lift over profile-only context+6.10 points

The registered participant-clustered contrast is not the subtraction of rounded descriptive means. Composite points summarize judge-rated response fidelity; they are not an approximation percentage, survey-distribution accuracy, or a percentage of correctly simulated individuals.

The positive generic comparison appeared in both corpora and under a second judge version. This supports the usefulness of relevant earlier evidence for these later interview questions, while leaving the strength of the effect in new domains open.

Why retrieval matters here

A compact profile captures enduring context, but a later question may require a particular earlier experience or reason. Retrieval can bring that evidence into the response without requiring every recorded detail to influence every answer.

The result is consistent with that explanation. It does not establish that more memory always helps, or isolate a universal effect of any one retrieval technique. The evaluated workflow, permitted context and held-out questions define the claim.

Limits and the named-model follow-up

The evidence set is small and spans only two corpora. Some source interviews were translated into English, which complicates judgments about voice and exact wording. Scoring used automated judges, not independent human raters, and should not be represented as human-certified fidelity.

The named foundation-model comparison reused these same 22 people and 66 answers with a different candidate set and scoring comparison. It is a follow-up evaluation on the same evidence, not an independent population replication. Its composite scores should not be combined with this study's scores.

The PRISM Alignment comparison provides a separate participant-fit dataset and different task boundaries. The evidence overview shows how these qualitative results complement, rather than replace, survey-distribution evidence.

The compact profile approximated the production training and retrieval workflow; it did not reproduce it exactly.

Reference data

De-identified, English Transcripts of 36 interviews

Transcription Data

CC BY 4.0