Research Methods September 10, 2026 · 7 min read

Synthetic Respondents Invent Customers Who Don't Exist

The 2026 validation research is in, and the thing AI panels are worst at is the exact thing you'd buy one for.

Two faceless retail mannequins against a grey background, one dark and one pale, standing in for survey respondents shaped like people but with nothing behind them
Edu

Edu

Founder, Insightios · About

Key Takeaways

  • In a segment-targeting test, AI panels inflated the gaps between segments two to fourfold and pointed at the wrong segment in half of the US cases (Chen, Zhu and Zheng, July 2026)
  • The same benchmark found the failures didn't shrink as models got bigger, tested from 8B parameters up to frontier scale
  • A psychometric audit of 37 models found they resemble each other more than they resemble people. Mean model-to-model similarity 0.733, best model-to-human 0.714
  • A plain statistical method with no AI in it beat every model at preserving how answers relate to each other, 0.95 against 0.52
  • Across twelve published experiments one review counted nine encouraging results against fourteen discouraging ones, and only 21% of classic studies replicated

Edu here. Someone asked me last month whether synthetic respondents were going to put me out of business, and I gave a lazy answer. I said the model has never bought anything. Which is true, and is also the kind of line you reach for when you haven't actually looked.

So I went and looked. The first thing I noticed had nothing to do with the technology: almost every page explaining synthetic respondents is published by a company that sells them (eight of the top ten results when I searched, most of them titled some version of "the 2026 guide"). Not a scandal, just how a new category gets written about at the start. But it does mean the buyer's half of this is mostly unwritten, and the research half turns out to say something fairly specific.

A July 2026 cross-domain benchmark of four large language models found that synthetic survey respondents inflate the gaps between customer segments by two to fourfold, direct a team to the wrong segment in half of US cases, and manufacture segment splits that do not exist in real people. The failures did not diminish as model capability increased.

Averages survive. Segments don't.

In 2026, researchers benchmarking four language models across two model families found that synthetic respondents inflate between-segment gaps two to fourfold, and would send a team after the wrong segment in half of US cases and most cross-cultural ones (Chen, Zhu and Zheng, When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses, July 2026). The models also produced segment splits that aren't there in the human data at all.

That's the expensive one. Nobody commissions a panel to learn what the average person thinks, because the average person isn't a customer. You commission it to find the group that behaves differently from the rest, and that specific output is the one the benchmark says you can trust least.

There's a second finding in the same paper that lands harder on DTC than on academia. The models treat demographics as far more predictive of attitudes than demographics actually are among real people. Think about how a DTC persona is usually built. Woman, 34, suburban, health-conscious, shops on mobile. Almost all of that is demographic scaffolding, and it's exactly the input the models over-read.

Then the part that should bother anyone waiting this out: the paper tested from 8B parameters up to frontier scale, and reported that neither failure improved with capability. So the usual reassurance (that the next release fixes it) has no support in the one benchmark that checked.

It isn't one paper

In April 2026, Jim Lewis and Jeff Sauro reviewed twelve published experiments with synthetic users and counted nine encouraging findings against fourteen discouraging ones (MeasuringU, A Review of Experiments with Synthetic Users, April 14, 2026). Their summary of the pattern is three words long and better than anything I'd write: superficial agreement, deeper errors.

The individual results underneath are worse than the tally suggests. Park and colleagues managed to replicate 21% of classic study findings with synthetic participants, so 79% didn't come back. Shrestha and colleagues found roughly 70% of 43 policy questions produced significantly different answers from humans. Almeida and colleagues found that even where correlation with humans was high, the models exaggerated effects by shrinking variance.

Only about a fifth of classic findings came back Only about a fifth of classic findings came back Replication of classic study results using synthetic participants (Park et al.) 21% replicated Replicated successfully 21% of classic findings Did not replicate 79% of classic findings Source: Park et al. (2024), as reported in Lewis and Sauro, MeasuringU, April 2026.

And the one closest to what a brand would actually run: Bisbee and colleagues found synthetic samples matched high-level means while producing inaccurate subgroup means, standard deviations too small, and inaccurate regression coefficients.

The models agree with each other more than they agree with you

A 2026 psychometric audit ran 37 models against a Lithuanian organizational psychology dataset of 263 employees answering 68 items, then scored how well each model preserved the statistical properties of the human responses (Lukauskas and Šarkauskaitė, Plausible but Not Valid: A Psychometric Audit of LLMs as Synthetic Survey Respondents, arXiv:2608.14606, 2026). Human-to-human similarity sets the ceiling at 0.825. The best model reached 0.714.

A Gaussian copula reached 0.688, a statistical technique with no language model in it, no persona, and no idea what a person is. It lands within three hundredths of the best LLM in the study. On preserving how answers relate to each other it isn't close at all: 0.95 for the copula against 0.52 for the best model.

Models resemble each other more than they resemble people Models resemble each other more than they resemble people Psychometric Similarity Score. Axis starts at 0.60 so the gaps are readable. Human vs human (ceiling) 0.825 Model vs model 0.733 Best model vs human 0.714 Copula baseline vs human 0.688 Source: Lukauskas and Šarkauskaitė, Plausible but Not Valid, arXiv:2608.14606, 2026. 37 models, n=263 human respondents.

The line I keep coming back to is the third bar. Mean similarity between one model and another was 0.733, higher than any model's similarity to actual humans. So they converge on each other. Running your study twice on two different models (which sounds like diligence) mostly buys you the same synthetic person wearing a different label.

I went in expecting the answer to be "it depends which model," and it mostly doesn't. That surprised me more than the headline numbers did.

The audit also found the models fabricated significant indirect effects on 3 of 10 null mediation paths, and produced gender effects against a human contrast where there weren't any.

Google's own team says the personas collapse

In February 2026, a team at Google DeepMind reported that even when a model is explicitly instructed to generate diverse personas, the output collapses around a narrow cluster of stereotypical responses (Paglieri et al., Persona Generators: Generating Diverse Synthetic Personas for Arbitrary Contexts, February 4, 2026). Asking for variety doesn't produce variety.

Their proposed fix is a multi-objective evolutionary search loop that deliberately hunts for rare attribute combinations and stress-tests the long tail. Real work, and it seems to help. Also nothing like what a brand gets from a subscription and a prompt box, and I think that distinction gets lost when vendors cite this line of research as validation.

Worth saying plainly: the people building the most sophisticated version of this technology published the clearest statement of its central weakness.

A hand using a stylus to tick a column of checkboxes on a tablet screen, representing survey questionnaire responses
The method being automated was already the weak link.

The part I skip when I make this argument

I have an obvious bias here (this is literally what I sell), so let me give the other side its best version. The survey was already leaking before AI touched it. People are bad at explaining their own behaviour to a stranger with a clipboard, they reconstruct reasons after the fact, and they answer the question they think you're asking. Synthetic respondents mostly automate a method that was already the weakest instrument in the room. If your alternative is a rushed 200-person panel written badly, I'm not sure the synthetic version is meaningfully worse.

Researchers seem to feel roughly this. Rival's 2026 Market Research Trends report, published December 4, 2025, put 42.75% of researchers as "not excited" about synthetic respondents (the page doesn't publish the sample or method behind that number, so treat it as a temperature reading rather than a finding). Not opposed. Not excited. I still think that's the most honest figure anyone has published on AI in research.

Here's what I can't tell you, and it's the thing I'd most want to know. Every paper above tests attitudes: social values, politics, workplace psychology. Nobody has run the same benchmark on "which of these three price points feels fair" or "why did you stop reordering." Maybe purchase questions hold up better, since they're more concrete and less identity-loaded. Maybe they're worse, because the stakes are imaginary and nothing was ever at risk. I genuinely don't know, and I'd rather say that than stretch a Lithuanian organizational psychology dataset into a claim about your skincare line.

Where the unprompted version lives

The reason I'm not that worried isn't that the models are bad. It's that the interesting answer was never the average one, and average is what a simulation is built to produce. When someone tells you why they stopped buying, the useful part is the weird specific reason: the scent changed, the cap leaked in a gym bag, their sister said it was a scam. A model gives you the most statistically ordinary version of a person, and ordinary people don't churn for ordinary reasons.

Those reasons are already written down, in public, by people nobody asked. In August I coded 3,100 comments about one haircare brand after an acquisition and found customers giving five incompatible explanations for the same complaint. A separate study of 8,000 comments about Liquid Death found 66% of people describing changed buying behaviour named the formula, taste, can size or price, while 8% named the marketing everyone assumed was the issue. No survey was involved in either. They were already arguing about it, for months, unprompted.

Messy, and it needs coding. But nobody had to imagine any of it.

Black and white photograph of a woman mid-gesture explaining something to a group in a crowded room, with other people listening around her
Nobody handed these people a questionnaire.

A simulated panel gives you the average customer. You don't have one of those.

Insightios goes to the threads where your market explains the problem to itself, codes what repeats, and hands back the phrasings, objections and buying triggers ranked by how common they are, with the real quotes attached. Flat fee, fixed turnaround.


Frequently asked questions

What are synthetic respondents?

AI-generated personas that answer surveys in place of people. You describe an audience, the system builds a panel of simulated respondents, and results come back in minutes instead of weeks. They're sold for concept testing and early pricing reads, usually as a faster substitute for recruiting a real sample.

Are synthetic respondents accurate?

They approximate high-level averages and fail underneath them. A 2026 MeasuringU review of twelve published experiments counted nine encouraging results against fourteen discouraging ones, and summarised the pattern as superficial agreement with deeper errors. One replication attempt reproduced only 21% of classic study findings.

Can synthetic respondents be used for customer segmentation?

Segmentation is where they perform worst. A July 2026 benchmark across four models found they inflate between-segment gaps two to fourfold, point at the wrong segment in half of US cases, and manufacture segment splits that don't exist in real people. Subgroup estimates are the least reliable output.

Will bigger AI models fix synthetic respondents?

There's no sign of it so far. The 2026 cross-domain benchmark tested models from 8B parameters up to frontier scale and reported that neither the individual-level failure nor the demographic over-determination improved with capability. The gap held across both model families tested.

Do AI personas produce diverse answers if you ask them to?

Not on their own. A February 2026 paper from Google DeepMind reported that even when models are explicitly instructed to generate diverse personas, the output collapses around a narrow cluster of stereotypical responses. Their proposed fix is an evolutionary search loop that forces coverage of rare trait combinations.

What can DTC brands use instead of synthetic respondents?

Comments people wrote without being asked. Reddit threads and reviews carry the specific reasons a purchase happened or stopped, in the customer's own phrasing. Insightios studies typically code between 3,000 and 8,000 comments per category to rank what actually repeats.


What to do with this

If you're already running synthetic panels, the practical read isn't "stop." It's that the outputs split into two piles of very different quality. Directional reads and questionnaire piloting sit in the safer pile. Anything shaped like a segment, a subgroup, or a demographic difference sits in the other one, and that's the pile most brands actually pay for.

Two questions I'd ask any vendor. What's the human baseline you validated against, and can I see it? And what happens to your numbers when you look at subgroups rather than the total sample? The second question is the one the published literature keeps answering badly.

Also, check the sourcing on whatever pro-synthetic statistic gets quoted at you. I chased several of the impressive ones circulating this year (a 90% correlation figure, an 85 to 95% distributional similarity range) and couldn't get any of them back to a named study with a disclosed method. The critical numbers in this post all come from papers with named authors, sample sizes and a methods section. That asymmetry isn't proof of anything by itself, but it tells you which side is currently showing its work.

All in all I don't think the question is whether simulated people are good enough yet. The problem is that they're built to produce the middle of a distribution, and everything I get paid to find lives at the edges of one. Maybe that changes. It hasn't yet, and if you want to know why someone stopped buying, they've probably already posted it somewhere. For a related test of a general AI research agent against a real corpus, see why deep research can't do voice of customer, and for the older version of this argument, why surveys fail DTC brands.


Sources

  1. Chen, Z., Zhu, D., & Zheng, L. N. (2026). When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses. arXiv:2607.26348, submitted July 28, 2026. Link Retrieved September 10, 2026.
  2. Lukauskas, M., & Šarkauskaitė, V. (2026). Plausible but Not Valid: A Psychometric Audit of LLMs as Synthetic Survey Respondents. arXiv:2608.14606. 37 models, Lithuanian organizational psychology dataset, n=263, 68 items. Link Retrieved September 10, 2026.
  3. Lewis, J., & Sauro, J. (2026). A Review of Experiments with Synthetic Users. MeasuringU, April 14, 2026. Review of twelve published experiments. Link Retrieved September 10, 2026.
  4. Paglieri, D., et al. (2026). Persona Generators: Generating Diverse Synthetic Personas for Arbitrary Contexts. Google DeepMind, arXiv:2602.03545, February 4, 2026. Link Retrieved September 10, 2026.
  5. Rival Group. (2025). 2026 Market Research Trends report, published December 4, 2025. Sample and methodology for the 42.75% figure are not disclosed on the report page. Link Retrieved September 10, 2026.
  6. Insightios research studies: Mielle Organics after the P&G acquisition (3,100+ comments) and Liquid Death reformulation loyalty (8,000+ comments). Analysis by the author.
Edu

Written by Edu

Founder of Insightios. I read Reddit threads, reviews, and support conversations so DTC brands can decide, price, and position from what customers actually say, not from what a survey guesses. More about me.