top of page

Real Results: Mixing Synthetic Data with Human Insights

  • Writer: Juan Jose (JJ) Ayala
    Juan Jose (JJ) Ayala
  • 4 days ago
  • 2 min read

I have been experimenting with synthetic data a lot in 2026, and it is fascinating. The research industry always drives for sample representation, usually focusing on the well-known "N=" number. Most researchers know that once you hit an N of 380, you sit comfortably in the 95% confidence interval with a 5% margin of error. From there, breaking data down into specific demographics shifts those confidence levels.


Recently, a survey showed that 69% of global researchers reported using synthetic responses in the past year. Using these synthetic customers alongside traditional research can yield comparable insights in half the time and at one-third of the cost. We wanted to figure out exactly how much we can trust these AI-generated proxies. Our main goals were to see if AI could pressure-test a questionnaire before it goes into the field, help narrow down a long list of ideas, and explore possible audience reactions early when recruitment is tough.


To answer these questions, I created a model that places real humans and synthetic respondents right next to each other. It involves a serious math model to create them, but once automated, it takes seconds. My biggest point of comfort is comparing the synthetic data with real humans in a strict 1:1 comparison table. This side-by-side benchmarking checks both the humans and the synthetics.


The numbers we are seeing tell a compelling story. In one academic study, researchers successfully reproduced 76% of original human main effects from previous marketing experiments using Large Language Model personas. Yet, 60% of researchers still oppose using synthetic samples, citing concerns about trust and authenticity. In the real world, major telecom providers are already testing features and pricing with synthetic customers to enter new market segments without damaging their premium brands.


I am launching a new product in the upcoming weeks for radio music programmers, giving you the best of both worlds. My model remains skeptical about synthetics standing alone and never leaves the real human data gathering and crunching behind. To make this dual approach work for your team, start by matching your research method directly to your objective.


Always recruit for relevance over sheer volume, because participants need the specific experience your questions depend on. Build a thoughtful incentive strategy that matches the effort required, and send that incentive quickly to build trust. You can also use AI to speed up analysis, organizing responses and clustering themes without outsourcing your own human judgment.


There are serious mistakes you must avoid. Research shows synthetic users have a pervasive "people-pleasing" tendency, frequently praising concepts without criticism. They might claim to easily finish tasks that would frustrate a real user to the point of giving up. Relying only on AI means you miss out on the messy, unexpected insights real people bring, like inventing a new workaround or using a product in a way you never planned.


Consumer research brings clarity to your biggest questions and confirms the path forward. By keeping real human benchmarks tied to synthetic speed, marketers can discover powerful, accurate answers without guessing.


Real Results: Mixing Synthetic Data with Human Insights

By Juan Jose (JJ) Ayala

Team Percepto, A Consumer Research Insights & AI Adoption Company, September 2026


Man compares AI and human traits on a glowing dashboard labeled Accuracy and Empathy, with charts, checkmarks, and handshake icons.

  • LinkedIn
  • Instagram
  • Facebook
  • TikTok
Team Percepto Logo

Copyright © 2024-2026 Team Percepto, LLC. All rights reserved.

bottom of page