---
title: "Personality Conditioning Makes Synthetic… | Minds"
canonical_url: "https://getminds.ai/research/personality-conditioned-synthetic-respondents-benchmark-2026"
last_updated: "2026-08-13T13:07:47.449Z"
meta:
  description: "In a blinded benchmark across 299 UK profiles and 869 held-out targets, personality-conditioned Minds were selected more often than generic or demographic..."
  "og:description": "In a blinded benchmark across 299 UK profiles and 869 held-out targets, personality-conditioned Minds were selected more often than generic or demographic..."
  "og:title": "Personality Conditioning Makes Synthetic… | Minds"
  "twitter:description": "In a blinded benchmark across 299 UK profiles and 869 held-out targets, personality-conditioned Minds were selected more often than generic or demographic..."
  "twitter:title": "Personality Conditioning Makes Synthetic… | Minds"
---

Minds

August 8, 2026·Validation·Minds Team

# **Personality Conditioning Makes Synthetic Respondents More Distinctive**

Minds achieved a 32.57% winner rate in a blinded three-arm comparison, with its clearest gains in values, specificity, and between-person differentiation.

[Read the Minds methodology](https://getminds.ai/research/methodology)

Personality-conditioned Minds were selected in **32.57%** of blinded, length-matched comparisons, ahead of demographic prompting at **25.66%** and generic prompting at **20.48%**; **21.29%** were ties.

The clearest advantage was not simply sounding human. It was sounding like a **particular** human: expressing values, adding relevant specificity, and producing more differentiation between people. That is the quality synthetic research needs when the goal is to explore how different audience members may frame the same question.

**32.57**%

Minds winner rate

**869**

Observed held-out targets

**1**

Blinded judges per item

Based on a simulated Audience of 299 respondents. Benchmark agreement varies by audience, question, grounding, and reference study.

## Why personality conditioning matters

Generic models are useful for producing a plausible answer. Market research usually needs something more demanding: answers that reflect meaningful differences in priorities, preferences, and worldview.

This benchmark asks whether a full Mind—grounded in demographics plus allow-listed, self-written personality evidence—can reproduce held-out qualitative expression better than the same model with no participant context or demographics alone.

The target was a participant-authored conversation opening, not a survey distribution or an objective “true” opinion. The result therefore measures qualitative profile fit: how well the response reflects the person-specific evidence available to the system.

## A blinded, three-arm benchmark

We used **299 UK participant profiles** from the [PRISM Alignment Dataset](https://github.com/HannahKirk/prism-alignment), pinned at revision `a9ec82149036374768f70e132a2369a750e06bbe`. The scored set contained **869 observed held-out targets**.

| Arm | Participant context | Execution |
| --- | --- | --- |
| Generic | No participant context | Direct Gemini 3.6 Flash |
| Demographic | Age, gender, country, employment, and education | Direct Gemini 3.6 Flash |
| Minds | Demographics plus allow-listed values, preferences, desired-assistant behavior, and training prompts | Production Panels API with persona, RAG, and web grounding enabled; Gemini 3.6 Flash |

Targets and target-derived data were prohibited from generation context. Each participant-domain output was generated independently. A single blinded Gemini 3.6 Flash judge saw the profile evidence, held-out human reference, and three anonymous candidates.

To ensure verbosity could not decide the result, every candidate was prefix-truncated to the human reference's word count before judging. All 869 targets received a complete three-arm judgment.

## Minds was selected most often

| Outcome | Generic | Demographic | Minds | Tie |
| --- | ---: | ---: | ---: | ---: |
| Winner count | 178 | 223 | **283** | 185 |
| Share of 869 items | 20.48% | 25.66% | **32.57%** | 21.29% |
| Judge-fidelity composite, 0–100 | 27.23 | 27.68 | **32.38** | — |

Minds' winner-rate advantage was **+12.04 percentage points** over generic prompting (95% CI +7.13 to +17.00; FDR-adjusted `p < .001`) and **+6.69 points** over demographic prompting (95% CI +1.56 to +11.93; FDR-adjusted `p = .020`).

The judge-fidelity composite improved by 5.10 points over generic prompting and 4.58 points over demographic prompting. Both differences survived multiplicity correction.

## The strongest signal: values and specificity

| Question context | Generic winner rate | Demographic winner rate | Minds winner rate |
| --- | ---: | ---: | ---: |
| Values-guided | 19.30% | 11.58% | **42.81%** |
| Controversy-guided | 18.88% | 34.97% | **38.81%** |

Minds reached **36.88** on values/personality and **33.46** on specificity, compared with 30.85 and 21.70 for generic prompting, and 31.85 and 23.36 for demographic prompting. Its genericness score was also lower—**63.63**, where lower is better—than generic (**73.53**) or demographic (**72.58**) responses.

The output was more differentiated between people as well. Mean between-person TF–IDF similarity was 0.0308 for Minds, compared with 0.0516 for demographic and 0.1108 for generic responses. Human references were more diverse still at 0.0128, but personality conditioning moved the model substantially closer to that pattern.

## A useful cohort signal

The most consistent improvements appeared in older cohorts. Ages 55–64 beat both baselines on winner rate and the composite after correction. Ages 65+ clearly beat generic prompting, while ages 45–54 improved over demographic prompting on the composite.

These subgroup results remain exploratory, but they suggest that richer accumulated preferences and values may give personality conditioning more useful signal to work with.

## Richer answers, fairly compared

Full Minds answers averaged **51.62 words**, compared with **7.86** for generic and **7.51** for demographic prompting. The benchmark's length-matching rule ensured that this added detail could not win simply by taking up more space.

That distinction matters commercially: Minds can generate richer material for qualitative exploration, while the next product step is to make response depth adaptive to the task. The artifacts do not contain comparable end-to-end token usage for every arm, so they do not support a monetary cost ratio.

## How to interpret the evidence

This is a strong directional result with clear boundaries:

- **One judge per item ( `n=1`).** The comparison used one blinded Gemini 3.6 Flash judgment and no human calibration.
- **One model family.** The response arms and judge ultimately used Gemini 3.6 Flash, which keeps the model constant but cannot rule out shared-model preference.
- **A qualitative target.** The benchmark tests fit to held-out human text, not population prevalence, purchase behavior, or lived experience.
- **An end-to-end system.** Persona conditioning, retrieval, grounding, and response formatting were evaluated together rather than as isolated components.

Semantic de-duplication was **not isolated in an ablation**, so this benchmark does not establish whether it helped or harmed the result and is not a basis for removing it.

We now use these 869 items to guide product development. Any future release-level claim requires a new, untouched holdout defined before evaluation, ideally with replicated judges or human calibration.

## What we test next

The next ablation will compare:

1. current production behavior;
2. responses constrained to human-like length; and
3. human-like length plus adaptive question routing that applies personality depth where it adds the most value.

This will show how much of the values, specificity, and differentiation advantage can be retained in a more adaptive response system.

## What this means for synthetic research

The practical finding is straightforward: **personality conditioning produced more distinctive qualitative respondents than generic or demographic prompting alone** in this blinded benchmark.

That makes Minds especially useful for generating and pressure-testing audience hypotheses, exploring different value frames, and surfacing the language different people may use. It complements rather than replaces human evidence—and gives research teams a more differentiated starting point than a generic model response.

## Suggested citation

Minds Research Lab. “Personality Conditioning Makes Synthetic Respondents More Distinctive.” August 8, 2026. https://getminds.ai/research/personality-conditioned-synthetic-respondents-benchmark-2026

## **Frequently asked questions**

### **What was the headline result?**

Personality-conditioned Minds were selected in 32.57% of blinded comparisons, ahead of demographic prompting at 25.66% and generic prompting at 20.48%; 21.29% were ties.

### **Where did personality conditioning add the most value?**

The clearest gains appeared in values expression, specificity, and differentiation between people. Minds won 42.81% of values-guided comparisons.

### **How was the comparison kept fair?**

Every arm used the same model family, candidate labels were hidden, and each candidate was length-matched to the held-out human reference before judging.

### **How should these results be interpreted?**

They show stronger qualitative profile fit in this benchmark. They do not by themselves establish population accuracy or universal human fidelity.