How accurate are synthetic respondents? What the evidence says in 2026
Vendors quote 85 to 95 percent parity. Independent reviews say accuracy has no universal number. Here is how to read the claims, what calibration actually means, and the one test that settles it for your business.
Published 15 September 2026 · 2 min read · By the Precheck team
Key takeaways
Published parity figures range from roughly 0.7 to 0.9 correlation with human responses, and vary by question type, audience, and model.
Parity on a benchmark is not accuracy on your decision. The only number that matters is how close the model got on decisions like yours.
Calibration means running the model blind against a known outcome and measuring the gap. Ask every vendor for it.
Precheck has not established predictive accuracy. We are running a calibration round to measure it rather than claim it.
Every synthetic research vendor has an accuracy number. Synthetic Users cites 85 to 92 percent synthetic-organic parity. Aaru cites a 0.90 median correlation in an EY engagement. Academic work on simulated survey responses reports model-question correlations from about 0.70 to 0.92 depending on the pairing. Independent reviewers land on a less quotable conclusion: there is no universal accuracy rate.
Both things are true. The numbers are real measurements. They just do not measure what a buyer needs to know.
Parity is not accuracy
A parity score compares a synthetic response set with a human response set on the same instrument. It answers: does the model produce answers that look like the answers people give? It does not answer: did the people then do what they said? Aaru's own positioning leans on this gap. Its argument is that humans are unreliable narrators of their own behaviour, so matching stated preferences is the wrong target.
For a pricing or bundle decision, the target is observed behaviour: what share actually bought at ₹1,299, whether the bundle cannibalised singles, whether churn moved after the packaging change. A model can score 0.9 against a survey and still miss those.
What moves accuracy up or down
From the published work and our own experience, four things dominate.
| Factor | Pushes accuracy up | Pushes it down |
|---|---|---|
| Underlying data | The company's own customer data for that audience | General-population training data only |
| Question type | Relative comparisons between scenarios | Absolute point estimates |
| Novelty | Decisions with close analogues in the data | Genuinely new products or categories |
| Validation | Blind runs against known outcomes in your category | Vendor benchmarks on public datasets |
The pattern is consistent with what NielsenIQ and Greenbook describe: accuracy is earned per audience and per task, and it decays as the question moves away from the data.
The test that settles it
There is one procedure that turns a claim into a measurement, and it is not complicated.
- Pick a decision you already made: a launch, a price change, a bundle.
- Give the model the inputs you had at the time, and nothing you learned afterwards.
- Let it simulate the response.
- Compare the simulated distribution with what actually happened.
- Repeat across a handful of decisions and audiences.
The result is a number that belongs to you and your category. If a vendor will not run this, or will only show it on their own examples, that tells you something about the number they quote.
Precheck's position
We have not established predictive accuracy, and our site says so on the page rather than in a footnote. Simulations are underway for a small number of large companies on data they have shared. We are inviting teams to bring a decision they have already made, with the outcome, so we can run it blind and share the comparison. If the simulation is wrong, we learn where. If it is close, that is the first honest accuracy claim we will make.
Questions people ask
- What does calibration mean in synthetic research?
- Calibration is running a simulation blind against a decision whose real outcome is already known, then measuring how far the simulated response distribution sat from what happened. It is the difference between a claim and a measurement.
- Are synthetic respondents accurate enough to replace surveys?
- For exploration and prioritisation, often yes. For decisions where the exact number matters, such as a price you will print on a shelf, no evidence supports skipping the real-world check.
- Why do accuracy figures vary so much between vendors?
- Because they are measured on different questions, audiences, and data. A model that tracks stated preferences well can miss actual purchase behaviour badly. Compare vendors on your decision, not their benchmark.