We ran Precheck blind against four real pricing experiments. Here is what happened.
Four published experiments where real customers were randomly shown different prices. We gave our simulation the situation, hid the answer, and compared. It ranked the prices correctly in eight of nine runs. It got the share of buyers badly wrong twice. Both results are below.
Published 20 September 2026 · 6 min read · By the Precheck team
Key takeaways
Across nine blind runs on four published experiments, the simulation put the prices in the right order eight times out of nine.
The share of customers who would buy was close in two cases and far off in two. On home-delivered water in Odisha it said 10 percent. Real households: 89 percent.
The misses have a pattern. The simulation repeats what experts already believe, so it fails where real customers surprised the experts.
What we will stand behind today: which price sells more, and how steeply demand falls. Not what share will buy. That needs real data from your category.
Every company selling simulated customers quotes an accuracy figure. Very few show you a case where the simulation was wrong. We think the misses are the useful part, so this write-up has both.
The test
We looked for published experiments where real customers were randomly shown different prices, and where the paper reports what share bought at each one. Random assignment matters. It means the difference in buying is caused by the price, not by who happened to see it.
We found four we could verify by reading the numbers in the paper ourselves. We rejected a fifth, on rainfall insurance in India, because its results by price appear only as a bar chart and we would not take them from a secondhand summary.
For each experiment we described the customers and the situation to Precheck, in plain words, without naming the company or the product's brand. We did not use web search to build the customer groups, because a search could have found the paper. Then we ran the simulation at the real prices and compared.
Two different AI models answer independently in every run. Nobody is ever asked what they would pay. Each simulated person sees one offer at one price and either takes it or walks away.
These were not Precheck customers. They are other people's experiments, used because the real answer is public.
Where it was close
Monthly subscription to an online job board
United States, 2015. 7,867 small businesses, each randomly shown one of ten prices at the paywall.
Average gap 6 percentage points. Two separate runs. Five prices first, then the five held-back prices with a newly written population. Real figures from Dubé & Misra, Journal of Political Economy, 2023 (Table 5).
This is the strongest result. Ten prices, from 19 to 399 dollars a month. The simulation put all ten in roughly the right place and was about six points under on average. We ran it in two halves on purpose: five prices first, then the five we had held back, with a completely new set of simulated people. The second half was as close as the first.
The paper's own headline was that demand was much less sensitive to price than the company assumed. Doubling the price from 99 to 199 dollars lost only a quarter of buyers. The simulation saw the same thing, a little more steeply.
Iron-fortified salt sold through village shops
Bihar, India, 2011. 4,179 households given vouchers at randomly assigned prices.
Average gap 9 percentage points. Only two price points are stated as numbers in the paper. The rest appear in a chart. Real figures from Banerjee, Barnhardt & Duflo, NBER (section 10.4.2).
An Indian consumer good, sold through village shops in Bihar. The simulation ran high this time, by about nine points, and got the size of the drop right: real buying fell by 82 percent between the two prices, simulated by 75 percent. The paper only states two prices as numbers, so this tests level and direction, not the shape of the curve.
Where it was badly wrong
Treated drinking water delivered to the door
Odisha, India. About 60,000 households in 120 villages, with discounts randomised by household.
Average gap 23 percentage points. Real figures from Burlig, Jina & Sudarshan, American Economic Review, 2026 (Table 1).
This is the result we most want you to see. At the lowest price, 89 percent of real households ordered clean water. The simulation said 10 percent. It missed the single most important fact in the experiment.
We read the simulated households' reasons. The arithmetic was right: about 42 rupees a month. They refused anyway, and nearly all gave the same reason. The handpump is free, so why pay.
That was also what most experts believed before this experiment. Earlier studies of water treatment had found low demand even when it was free. The researchers found the opposite: households valued clean water delivered to the door several times more than anyone had estimated. That surprise is why the paper was published where it was.
So the simulation did not fail at random. It reproduced the conventional view, and the conventional view was wrong.
Malaria bed nets offered at prenatal clinics
Western Kenya, 2007. 424 women across 16 clinics, with price randomised by clinic.
Average gap 37 percentage points. One of the two models could partly recall this paper's results when asked directly, so this is the weakest test of the four. Real figures from Cohen & Dupas, Quarterly Journal of Economics, 2010 (Table 3).
Here the simulation was far too pessimistic. A quarter of simulated pregnant women turned down a free malaria net offered by their nurse. In reality almost nobody did.
This one was our fault, and we found the cause. AI-simulated customers are known to overstate what they would pay, so we had instructed ours, firmly, that most people refuse most offers. That correction went too far. The simulated women invented reasons that were not in their profiles: a net they already owned, a relative who had to approve.
What we changed, and why it does not count
We rewrote the instructions so that the offer itself sets how likely people are to say yes, and so that simulated people may not invent obstacles they were not given. Then we ran all four again.
| Experiment | Average gap, first attempt | Average gap, after the change |
|---|---|---|
| Job board subscription, United States (the five prices run both times) | 5 points | 4 points |
| Iron-fortified salt, Bihar | 9 points | 15 points |
| Delivered water, Odisha | 23 points | 26 points |
| Bed nets, Kenya | 37 points | 15 points |
Bed nets improved a lot. Salt got worse. Water did not move at all, which confirms that the water miss is about what the models believe, not how we asked.
We are not claiming the second column as a result. We changed the engine after seeing these answers, so these four cases can no longer tell us whether the change is a real improvement. Only experiments we add afterwards can. The charts above all show the first attempt for that reason.
What this means if you are deciding a price
The order of prices is reliable. In eight of nine runs the simulation ranked every price correctly, and in the ninth it was close. If you want to know which of three prices will sell most, and roughly how fast buying falls as the price rises, the simulation is useful today.
The share who will buy is not reliable on its own. It was within ten points in two categories and off by as much as eighty in another. A simulation that tells you "27 percent will buy" is giving you a number it cannot yet support.
It cannot find a surprise. The simulation knows what is broadly known. If your customers value something more than the market assumes, which is often exactly why a new product works, it will miss that. Real data from your own customers is what fixes this. When you share order history, the customer groups are built from your numbers rather than from the model's assumptions.
The consistency checks are not truth checks. Every run is checked for whether the two models agree and whether demand falls as price rises. Both of the bad misses passed those checks. They tell you the simulation is coherent, not that it is right.
This is why Precheck ends every run by recommending one real-world test rather than a final answer. The simulation narrows three prices to the one worth testing. The test tells you the truth.
What happens next
We are adding more published experiments, with a preference for Indian consumer categories, and we will keep the results here whether they flatter us or not.
The more valuable test is yours. If you have already made a pricing, bundle or packaging decision and know how it turned out, send it to us. We run it blind, and you get the comparison.
Sources
- Dubé, J-P. and Misra, S. "Personalized Pricing and Consumer Welfare." Journal of Political Economy, 2023. NBER working paper 23775.
- Banerjee, A., Barnhardt, S. and Duflo, E. "Nutrition, Iron Deficiency Anemia, and the Demand for Iron-Fortified Salt: Evidence from an Experiment in Rural Bihar." NBER chapter 12984.
- Burlig, F., Jina, A. and Sudarshan, A. "The Value of Clean Water: Experimental Evidence from Rural India." American Economic Review, 2026. NBER working paper 33557.
- Cohen, J. and Dupas, P. "Free Distribution or Cost-Sharing? Evidence from a Randomized Malaria Prevention Experiment." Quarterly Journal of Economics, 2010. NBER working paper 14406.
For the water and bed net studies, the real figures are the comparison group's rate plus the estimated effect of each price, as reported in the papers' regression tables. They are close approximations of the raw rates rather than exact cell averages.
Questions people ask
- Were these Precheck customers?
- No. These are published academic experiments run by other people, years ago, on their own customers. We used them because the real outcome at each price is public, which makes them a fair test. Precheck had no part in the original studies.
- Could the AI models have simply remembered the papers?
- We checked. We named each paper to both models and asked them to recall its results. Neither could for three of the four. One model partly recalled the bed net study, so we treat that case as the weakest. The company or product was never named to the simulation itself.
- Is this an accuracy claim?
- No. Four experiments is a first look, not a track record, and we changed the engine after seeing these results, so any later score on these same cases does not count. Predictive accuracy has not been established. This is the start of measuring it in public.
- Can I see how it does on my own past decision?
- Yes, and that is the most useful test there is. Send a decision you already made, the options and prices you weighed, who it went to, and what happened. We run it blind and show you the comparison, hits and misses alike.