Knowsis Logo
KNOWSIS
Decision intelligence for insights teams
Interactive Tool

The Synthetic Data Ladder

Not all synthetic data is the same thing. It runs from harmless placeholder numbers up to models that answer as a specific person. Run this two-minute check to see which rung you are on, what to ask, and whether to trust it for your decision.

The rule: the bigger the claim, the harder the test.   The default: synthetic for direction, humans for decision.
Step 1

Which rung is your data on?

Pick the one that best describes what you have been handed.

Step 2

What is riding on the decision?

Be honest about the cost of being wrong.

Step 3

Five quick checks

Answer what you can. "Not sure" is a real and useful answer.

Is there a real person underneath it? (built from real respondents, not a model's general memory)
Has it been tested against real outcomes or held-out real data, not just ‘does it look right’?
Is it current for the market you actually care about?
Does it keep the real differences and relationships, rather than flattening everyone toward the average?
Are you using it for the average and the direction, not to predict an individual or an exact number?
Your read
Pick your rung and your stakes above, and your read will appear here.
The research worth knowing

Durable evidence, not the daily debate

Two studies and two regulators worth reading before you trust any synthetic dataset.

Study · concept testing
Semantic Similarity Rating (SSR)
PyMC Labs & Colgate-Palmolive, 2025
LLM ‘synthetic consumers’ matched about 90% of what real surveys found on concept-testing questions, and recovered the ranking of concepts, by reading their worded reactions rather than asking for a number.
Read it → arXiv:2510.08338
Dataset · digital twins
Twin-2K-500
Toubia et al., Columbia Business School (Marketing Science), 2025
A public dataset of 2,000+ real people answering 500 questions, built as ground truth for creating and testing digital twins of individuals. The benchmark that lets you check twins honestly.
Read it → arXiv:2505.17479
Regulator · privacy & validation
Synthetic data: validation and privacy
UK Financial Conduct Authority
Frames quality as fidelity, utility and privacy, and warns that synthetic is not automatically anonymous. Regulator-grade, not vendor marketing.
Read it → fca.org.uk
Regulator · governance
Proposed Guide on Synthetic Data Generation
Singapore PDPC
Practical guidance on generating and governing synthetic data, including re-identification risk. A clean checklist for provenance and disclosure.
Read it → pdpc.gov.sg

Want to see this test run on personas?

We put our own predictive personas through the hardest version of these checks: built on part of a public sample, tested on people the model had never seen, with the method written down first. Read how they held up, or bring your own data to a call.

Book a call