Asking a vendor to show you their validation gets you a validation deck, and a validation deck is a document written by the person selling to you. These are the questions that actually check the number.
None of these questions needs a statistician in the room, and none of them is hostile. A supplier with a sound product can answer all eight in an afternoon, because they will have had to answer them for themselves before they could trust their own output.
Each one carries a number from public validation work showing why it matters. Where the number is ours, it is not flattering, which is rather the point.
Every accuracy figure has a floor underneath it: the score you would get knowing nothing about the individual at all. Almost nobody reports it, and it is usually most of the number.
On Columbia's public Twin-2K-500 benchmark we scored a single answer per question, the same for every respondent. It reached 0.698. The best digital twin in the published study, built from around five hundred of each person's own answers, reached 0.712.
Bars run from zero, so the picture is honest: the gap between guessing and knowing nothing is large, and the gap between knowing nothing and knowing five hundred things is 1.4 accuracy points.
The same metric, run on a no-information baseline, on the same people and the same questions, shown next to the headline figure. If they have never computed it, that is your answer.
A vendor takes their accuracy, divides it by how consistently real people agree with themselves when asked the same question twice, and reports the result as a normalised or calibrated score. Nino Hardt summed it up better than we could: divide by 0.8 and call it normalised in an asterisk.
Correcting for an unreliable benchmark is a century old and perfectly legitimate idea. The problem is that simple division never subtracts the floor, so everything gets credit for the base rate.
Here is what that does, on the same benchmark and the same people. Divided by human retest, knowing nothing about a person scores 85%. Demographic segments score 85%. Our driver-based personas score 86%. The best twin scores 87%. Everything looks excellent and nothing looks different.
Measure the same four results as a share of the distance from random guessing to the human ceiling and the picture changes: 53%, 54%, 57% and 59%. Still tightly bunched, because the differences really are small, but now honestly placed in the middle of the range rather than near the top of it.
How the figure was normalised, and what chance scores on the same scale. A vendor quoting 85% normalised may be quoting the base rate.
A model tested on the people it was built from is marking its own homework. It will always look better than it is, and the gap is not small.
We published this against ourselves. Our persona method separates the people it was built from at 0.61. On people it had never seen, it separates them at 0.20. The honest number is a third of the flattering one, and the flattering one is the one that is easy to quote by accident.
The figure comes from held-out respondents who played no part in building or tuning the model, and they can tell you how the split was made. "We validated it on our panel" is not an answer until you know whether the panel was also the training data.
A single overall figure is an average over people the method understands and people it does not. The average can hide almost anything.
Across the subgroups in our test, accuracy ran from 0.662 to 0.726. That spread is more than four times the width of the entire gap between knowing nothing and the best twin. One of our personas scored above the twin's overall figure. Another scored below a single answer given to everybody.
Columbia's published work found the same shape in its twins, which were more accurate for more educated, higher income participants. Independent work on the World Values Survey found synthetic accuracy holding up for Western, English-speaking, wealthy samples and degrading elsewhere. If your market is not one of those, ask this question twice.
Results broken out by the segments your decision actually depends on. If the vendor only has an overall figure, they do not yet know who their product is wrong about.
Two thousand synthetic respondents produce a table that looks like a two thousand person survey, with confidence intervals and significance tests to match. Those intervals are almost always calculated as though each synthetic respondent were an independent person.
They are not. Every one of them comes from the same model, so their errors move together. When one is wrong in a particular direction, the others tend to be wrong the same way. The real amount of independent information is far smaller than the row count, and the intervals are correspondingly too narrow.
This is not only our argument. Ipsos's own data science team made it in print this year: synthetic records cannot carry the statistical confidence of real ones, and the effective sample size levels off however many rows you generate. Separately, Prolific found that simulating respondents one at a time produced roughly three times the error of simply asking the model for the aggregate distribution.
An effective sample size, or a straight statement that the intervals were computed from real data rather than from the generated rows.
This is the question that separates a product from a claim, and it is where the field is quietly converging.
Columbia's published paper found twins less variable than the real people they modelled in 93.9% of outcomes. FlashInsight, unusually, published the diagnosis of their own engine: across eight markets their synthetic respondents disagreed with each other at between 0.20 and 0.76 of real human levels before correction, and between 0.96 and 1.07 after it. They also publish what they have not yet done, which is the out-of-sample test on new questions. A separate research paper this month showed that a calibration sample of between 50 and 300 real respondents removed most of the bias in population-level estimates, and validated it on the same Columbia dataset.
Notice what those corrections do and do not fix. They make the group-level picture right: the averages, the spread, the proportions. None of them claims to make any individual prediction better.
Once there is a real sample, one figure falls out of it that is worth more than any accuracy claim. How well does the synthetic output correlate with the real answers, on the outcome you actually care about? That single number sets a ceiling on everything the product can be worth, and no amount of method can lift it.
At a correlation of 0.4, a synthetic prediction can be worth about 16% more sample. At 0.2, about 4%. At 0.06, it is worth nothing at all, and the cost of working out how much to trust it eats more than it returns. When we measured our own, the personas correlated at 0.392 with the outcome they were built around and delivered 14%. On individual questions they correlated at 0.063 and delivered nothing.
A random sample of real respondents the synthetic output is checked against, how large it is, how often it is refreshed, the before-and-after numbers, and the correlation with real answers on your outcome. If there is no real sample anywhere in the process, there is nothing to calibrate against, and the confidence being offered has not been earned.
A validity claim has to name its level. A persona is a group claim, judged on whether groups genuinely differ. A twin is an individual claim, judged on whether it gets one named person right. The same technology underneath, and a completely different burden of proof. Most of the public argument about synthetic respondents is two people measuring different things and disagreeing about the result.
It also has to name its outcome, and that half is easy to miss. When we priced our own synthetic predictions against real respondents, the personas were worth about one extra real respondent in seven on the outcome they were built to explain. On the other 125 questions in the same study, they were worth nothing, and in small samples slightly less than nothing. Same personas, same people, same run.
A claim that names both: this is accurate at group level, on this outcome. Be wary of anything validated on one outcome and sold for all of them.
Two different failures sit inside this one question, and both are easy to walk into.
A corpus has a date. Ask a model the same question twice and you get its variance, not the market's, so a synthetic tracker measures the model rather than the movement you are trying to detect. If you want to use one for change, the test is whether it has reproduced a past shift you already know the answer to.
And outside the training set there is nothing to interpolate from. What comes back is a plausible prior with a confidence interval attached, which is the most dangerous possible combination, because it looks exactly like a measurement.
An honest statement of what the output is for. Screening and direction inside the territory the model knows is a reasonable claim. Tracking change, or answering about something genuinely new, is not, unless they can show a back-test on a shift that already happened.
The first six came out of a series testing synthetic respondents on public data, including our own method. Questions seven and eight were added in public by people who read it: Sandeep Budhiraja of Spark Sensory AI, and JD Deitch of YouGov. Both improved the list, and both are credited above.
If you think there is a ninth, we would like to hear it. The list has grown twice already.
They are not flattering. Our method scores a third of its in-sample figure on strangers, it does worse than a constant for one of its four groups, and on an individual accuracy scale it sits inside a band barely wider than the noise. We published all of that, because the alternative is asking buyers to trust a number we would not trust ourselves. The full working is in The Broken Ruler Eight Questions Checklist.
Used for direction, screening and ordering, inside the territory the model knows, it earns its place. The questions are about whether a specific number has been checked, not about whether the technology works.
Eight good answers tell you a supplier has done the work. They do not tell you the product will work on your category, in your market, on your decision.
It was six three weeks ago. The field is moving quickly enough that a checklist which does not change is a checklist that has stopped paying attention.
The validation study is public, the method was fixed in writing before each test, and the figures come from data anyone can download. Or bring a vendor's deck to a call and we will go through it with you.