We tested our predictive personas on a public dataset of 2,058 people, the same dataset behind the study most often held up as proof that AI personas fail. The personas held. We also found out why honest people keep reaching opposite conclusions about synthetic respondents.
Knowsis builds predictive personas from survey data. A client picks a persona and interrogates it. Our credentials pack has always promised three specific validation checks, but promising a check is not the same as having run one.
Columbia Business School gave the whole industry a way to fix that. They published a free dataset of 2,058 Americans who each answered around 500 questions, then used it to build AI copies of individual people. Their follow-up study found those copies performed no better than a cheap stand-in built from age, income and location. The field read that as proof that AI personas do not work.
We think that conclusion tests the wrong claim. So we ran ours on the same data.
A digital twin is an individual claim. It says "this is Bob", and you judge it on whether it gets Bob right. That is what Columbia tested, and that is what came back null.
A persona is a group claim. It says "people like this tend to behave like this", and you judge it on whether the groups genuinely differ. Same technology underneath, completely different burden of proof.
We picked one outcome to predict: 40 questions where each person saw a real grocery product at a real price and said whether they would buy it. Butter at $7.39, cat food at $1.58. Each person's answers became a single purchase score.
Then we split the sample. Everything was built on 80% of respondents, and the other 20% were held back entirely. The only test that counts is whether personas built on one group say anything true about people the model has never seen.
Before each stage we wrote down the method, the comparison and the pass mark, then ran it. That sounds like bureaucracy. It is the difference between running a test and staging a demonstration.
Both approaches got the same information and the same hidden respondents to sort. Personas built on what drives behaviour explained roughly three times as much of the difference between people as segments built from age, income, region and the rest.
Driver-based personas recover three quarters of what any four-group split of the best available model could achieve. Demographic segments sit barely above what random grouping produces.
The obvious challenge is that our personas were allowed to use demographics too, so perhaps they simply repackaged them. We rebuilt from the psychological measures alone. They still beat demographic segments roughly two to one, and putting demographics back improved things by an amount we could not distinguish from zero.
On this sample the demographic information was actually more predictive of an individual's score. It still produced half the separation once you group people with it. Same information, organised differently, twice the result. The difference is not what you know about people. It is whether your segments are built on modelled drivers of the outcome or on the raw variables.
Holding out people is a good test. It is not the hardest one, because both groups answered in the same sitting, so any quirk of that day sits in the training data and the test data alike.
Columbia re-asked the same questions two weeks later. So we predicted the second sitting using personas built on the first, for people the model had never seen. Separation came back at 102% of the original figure. Nothing was lost.
Build personas now and the group differences they describe are still there a fortnight later, at full strength. For a tracking or continuous programme, that is the property that matters.
Every version we built put the personas in the right order on new people. None of them got the levels right. A persona scoring 66 on the group it was built from scored 60 on strangers.
Comparing the two sittings told us why, and it is good news. The compression happens when you move to new people, not when you move forward in time. That makes it a correctable bias rather than drift.
Use it to decide which option is better, which segment to prioritise, which direction to move. Do not quote its score as a forecast. Ask a persona about one specific product and it answers weakly. Ask it to order a basket, a category or a set of concepts and it answers three and a half times better.
The dataset lets you ask something almost nobody publishes: how often does a real person give the same answer they gave a fortnight ago?
Across the repeated questions, individuals repeated themselves 69% of the time. Ask instead whether the overall percentages held between the two rounds, and the answer is 97%.
That 28-point gap explains most of the public disagreement about whether any of this works. A vendor quoting the first number and a critic quoting the second can both be telling the truth, and they will never agree, because they are not discussing the same quantity.
It is also why the most impressive-sounding figures in this category are the ones to check first. When someone reports an accuracy number, the useful question is not whether it is high. It is what it was measured against, and at what level.
American respondents, grocery purchase intent. Nothing here speaks to other markets or other kinds of decision.
People's answers barely shifted by product or price. Two items were priced at zero and still drew a normal response rate. The measure is closer to general willingness to buy than to price sensitivity.
Four personas divide a fairly continuous spread of people rather than uncovering four natural clusters. Normal for segmentation, and worth saying out loud.
This ran without our motivational layer, which needs a question battery this dataset does not contain. The result was achieved with part of the toolkit. The two-minute quiz is that layer, run on you rather than on a sample.
That has been our position since before we ran any of this, and the validation supports it rather than a stronger version of it. Personas built on measured drivers tell you reliably which groups differ and in which direction. They do not tell you what a number will be, and anyone claiming otherwise should be asked how they measured it.
If you commission this kind of work, three questions are worth asking of any supplier, including us. What does the figure score when you take the individual out of it? Was it earned on a group-level task or an individual-level one? And can you recompute it from what has been published?
What was tested here is not a one-off. It is the process we use on commercial work, and the argument behind it was published in Greenbook in September.
The Predictive Persona Playbook is the process this validation was run on. The Synthetic Data Ladder is a two-minute check on whether anyone's synthetic data is worth trusting for your decision. The Impulse Engine quiz is the motivational layer this test ran without, put to you rather than to a sample.