---
title: "Model-Change and Validation Provenance Policy | Minds"
canonical_url: "https://getminds.ai/research/model-change-validation-provenance"
last_updated: "2026-08-26T19:29:45.503Z"
meta:
  description: "How Minds records material model, provider, prompt, retrieval, population, orchestration, post-processing, and calculator changes that affect research claims."
  "og:description": "How Minds records material model, provider, prompt, retrieval, population, orchestration, post-processing, and calculator changes that affect research claims."
  "og:title": "Model-Change and Validation Provenance Policy | Minds"
  "twitter:description": "How Minds records material model, provider, prompt, retrieval, population, orchestration, post-processing, and calculator changes that affect research claims."
  "twitter:title": "Model-Change and Validation Provenance Policy | Minds"
---

Minds

August 21, 2026·Governance·Minds Team

# **Model-Change and Validation Provenance Policy**

Minds will not carry a public validation claim across a material research-system change without preserving the tested system envelope, assessing impact, and rerunning scoped checks where comparability may change.

[Review Minds evidence](https://getminds.ai/research/synthetic-research-evidence-center)

Synthetic research systems change continuously. A new frontier model can improve instruction following while changing response distributions. A retrieval update can add better evidence while altering which facts an audience member sees. A calculator fix can improve an estimate without changing a single answer. A public validation claim must therefore describe the tested system envelope, not float above the product as a timeless percentage.

Minds commits to preserving the provenance of public validation claims and to assessing material research-system changes before treating earlier results as directly comparable.

## The commitment

For a public Minds benchmark or validation result, we will:

1. date the study and identify the population, instrument, reference, and metrics;
2. describe the material system layers relevant to interpretation;
3. retain the result definition and known limitations;
4. assess whether a later material change can affect the claim;
5. run scoped regression or benchmark checks when comparability may change; and
6. disclose whether a later result is a reproduction, extension, or new benchmark rather than silently replacing the old result.

This is a disclosure and evaluation commitment. It is not a promise that every output remains identical after every model update.

## Changes in scope

The review covers changes to:

- base model or model version;
- model provider or inference region;
- system prompts and answer contracts;
- retrieval, embeddings, ranking, or source policies;
- audience construction, distribution allocation, or persona profiles;
- orchestration, agent interaction, and post-processing;
- deterministic method calculators and diagnostics;
- benchmark data, metric definition, or analysis code; and
- safety or refusal behavior that changes study completion.

Customer-provided source updates and deliberate study-configuration changes also affect comparability, even when the product system is unchanged.

## Impact levels

| Level | Example | Expected action |
| --- | --- | --- |
| No research impact | Styling or unrelated interface copy | No benchmark action |
| Low | Presentation change that preserves stored answers and estimates | Confirm artifact parity |
| Moderate | Prompt, retrieval, or classification change affecting some answers | Run focused regression on affected tasks |
| High | Base-model, provider, population, estimator, or metric change | Rerun the relevant benchmark or publish a new version |
| Data-flow impact | New provider, subprocessor, region, retention, or source policy | Security, legal, and procurement review plus notice where applicable |

Impact is determined by the claim at risk. A calculator change may be high impact for a conjoint result and irrelevant to a qualitative transcript benchmark.

## The tested system envelope

A public report should preserve enough context to interpret the result, which may include:

- study and execution date;
- audience identifier or reproducible definition;
- source-bundle and outcome-isolation description;
- method and calculator versions;
- model family or material provider information where disclosure is appropriate;
- prompt or answer-contract version identifiers;
- completion, fallback, and repair behavior;
- metric code or unambiguous formula; and
- limitations and known exposure risks.

Minds will not publish credentials, private customer content, security-sensitive internal details, or protected participant data to satisfy provenance. Reproducibility and privacy must both be designed into the artifact.

## Regression and benchmark rules

A regression check asks whether a known workflow changed unexpectedly. A benchmark asks how well the system agrees with an external reference. They are not the same.

After a moderate change, Minds may rerun a frozen internal regression set for the affected answer contract, method, subgroup, or language. After a high-impact change, a public claim should be re-evaluated on the relevant external benchmark before the earlier number is presented as current-system evidence.

If the original source is public, the report should continue to acknowledge possible indirect pretraining exposure. If the benchmark definition changes, the new result must not be shown as a simple improvement over the old score without a bridge analysis.

## Public benchmark versioning

When a public validation is rerun, the result should be labeled as one of:

- reproduction: same population definition, instrument, reference, and metrics;
- extension: adds populations, instruments, metrics, languages, or subgroups while preserving a comparable core;
- replacement: corrects a material error in the original analysis; or
- new benchmark: changes the research question or measurement enough that direct ranking is inappropriate.

Older reports should retain their date and limitation rather than being silently rewritten into the latest claim.

## Model-provider changes and procurement

A provider change can alter more than research performance. It may affect subprocessors, processing locations, retention, zero-data-retention settings, contractual terms, and incident dependencies. Those changes enter the procurement review separately from benchmark performance.

Minds publishes a [subprocessor list](https://getminds.ai/legal/subprocessors), [DPA](https://getminds.ai/legal/dataprivacy), [TOM](https://getminds.ai/legal/tom), [SLA](https://getminds.ai/legal/sla), and [DPIA](https://getminds.ai/legal/dsfa). Applicable notices and customer-specific terms remain governed by the relevant agreement.

## What buyers should record

For a consequential study, the buyer should retain:

1. decision and risk level;
2. audience and source versions;
3. instrument and stimulus versions;
4. method and calculator versions;
5. system and model envelope;
6. raw and derived artifacts;
7. benchmark or validation evidence;
8. approvals and required live-evidence step.

Use this policy with the [research evidence center](https://getminds.ai/research/synthetic-research-evidence-center), [validation and accuracy comparison](https://getminds.ai/comparison/synthetic-audience-validation-and-accuracy), [method pipeline catalog](https://getminds.ai/research/research-method-pipeline-catalog), and [procurement checklist](https://getminds.ai/guide/synthetic-research-procurement-checklist).

## **Frequently asked questions**

### **Why does Minds publish a model-change policy?**

Synthetic research quality depends on more than a model name. Provider APIs, prompts, retrieval, audience construction, orchestration, post-processing, calculators, and source data can all change results and benchmark comparability.

### **Does every software change require a full benchmark rerun?**

No. Changes are assessed by likely impact. Material changes that may affect a public claim, data flow, or result comparability trigger scoped regression or benchmark work; cosmetic and unrelated changes do not.

### **Will Minds disclose every internal prompt?**

No. Provenance requires enough information to interpret and reproduce the claim envelope without exposing security-sensitive implementation detail, private customer material, credentials, or protected intellectual property.