Skip to main content

Human Data Remain Essential in the Age of Synthetic Respondents

Expert View
A human in a metaverse: An AI evolving, creating unique meta avatars in a digital universe. Firewall against viruses, creating copies.

Author

Leah Christian

Senior Vice President

Methodology & Quantitative Social Sciences

September 2026

Artificial intelligence is transforming nearly every industry, and survey research is no exception.

Over the past two years, researchers have increasingly explored the use of AI-generated “synthetic respondents” to answer surveys, test ideas, and generate insights. These approaches have attracted significant attention and sparked vigorous debate across academia, government, and industry.

My first experience with synthetic respondents came nearly a decade ago while working on audience measurement solutions. We had integrated panel and device data but only had aggregate estimates across all the data sources, so we used machine learning models to generate respondent-level data that aligned with those broader patterns. Looking back, it was an early version of a challenge that has now become a major topic of discussion across survey research.

Fast forward to today, and the interest in synthetic respondents has grown enormously. It reminds me of many of the innovations that have emerged during my career, each accompanied by bold claims, understandable concerns, and important methodological questions. Whether it was online surveys, passive measurement, or new forms of digital data, the most useful innovations were not the ones that generated the most excitement. They were the ones that were subjected to careful testing and shown to be fit for purpose. That perspective shapes how I think about synthetic respondents today. They represent a promising area of research, but one that requires the same scientific rigor, transparency, and evaluation that we would expect of any new methodology.

“Synthetic respondents represent a promising area of research, but one that requires the same scientific rigor, transparency, and evaluation that we would expect of any new methodology.”

Senior Vice President, Methodology & Quantitative Social Sciences

“Synthetic respondents represent a promising area of research, but one that requires the same scientific rigor, transparency, and evaluation that we would expect of any new methodology.”

Early Interest & Challenges with Synthetic Approaches

Terms such as synthetic responses, synthetic panels, synthetic personas, synthetic respondents, digital twins, and silicon sampling are often used to describe related approaches: using AI models to generate responses that attempt to mimic how real people would answer survey questions.

The appeal is clear. Synthetic respondents can be generated quickly, at relatively low cost, and without adding burden to respondents. Researchers have proposed using them for everything from questionnaire testing and rapid prototyping to augmenting existing survey samples and exploring rare populations.

But this excitement has been matched by legitimate concerns. Early research has found that large language models do not answer survey questions in the same way people do, which raises questions about how accurately synthetic respondents can represent real individuals. Research has also found that AI-generated respondents may not reflect the full range of answers seen in survey data. For example, synthetic respondents may produce fewer very positive or very negative responses than are typically observed among human respondents, resulting in distributions that are more concentrated around the middle of the scale.

The relationships between answers to different questions may also differ from those found in human survey data. For instance, respondents generated by an AI model may provide more internally consistent answers across related topics than people do in practice, leading to patterns that look different from real-world survey results. Finally, results can vary depending on how information is provided to the model, including the amount of background information available and how respondent characteristics or survey questions are presented. Researchers have found that changing the context provided to the model, adding prior survey responses, or describing a respondent in different ways can produce different answers.

Together, these challenges raise important questions about accuracy, representation, inference, robustness, transparency, and the importance of careful evaluation. For me, these findings suggest that the most important question is not whether synthetic respondents are possible, but whether they can be developed, evaluated, and used responsibly for specific research purposes.

NORC’s Focus on Rigorous Evaluation & Experimentation

NORC has launched a systematic research program focused on understanding when, where, and how synthetic respondents may be useful. Our objective is not to find a one-size-fits-all solution, but to determine which approaches are fit for purpose and under what conditions.

Our research examines many of the methodological factors that prior literature suggests are important, including model selection, context engineering, supervised fine-tuning, and calibration. In practice, this means evaluating which AI models perform best, how the information provided to those models influences their responses, whether additional training on high-quality survey data improves performance, and how outputs can be adjusted to better reflect observed patterns in real populations. We also investigate broader data considerations such as sampling approaches, representativeness, historical respondent information, and the timeliness of underlying data sources.

Beyond modelling, our research focuses on aspects of the data including sampling frameworks, representation, longitudinal information, and data recency. Our work is built on NORC’s probability-based panel, AmeriSpeak®, because it provides known selection probabilities, rich respondent profiles, historical survey responses, and high-quality human-collected data that can serve as both training material and evaluation benchmarks. What excites me most about our work is that synthetic respondents are never created in isolation. They are anchored to real respondents and real survey data.

Our experimental framework currently evaluates dozens of combinations of models, prompting strategies, and contextual information, with plans to test hundreds more. We compare multiple AI models, different prompting designs, varying levels of respondent profile information, and a range of survey topics spanning health, employment, income, media consumption, and political attitudes. Importantly, we evaluate performance not only overall but across demographic subgroups and multiple dimensions of data quality.

“High-quality probability-based human data provides the foundation we need to build, evaluate, calibrate, and govern synthetic systems.”

Senior Vice President, Methodology & Quantitative Social Sciences

“High-quality probability-based human data provides the foundation we need to build, evaluate, calibrate, and govern synthetic systems.”

Human Data Are Not Optional

Every stage of the synthetic response process depends on high-quality human data. 

The most important lesson from our research so far is that human-collected data become more important, not less, in a world of synthetic respondents. Human data are needed to train models, calibrate outputs, evaluate accuracy, identify bias, assess subgroup representation, and determine whether synthetic respondents are fit for purpose. They provide the foundation we need to build, evaluate, calibrate, and govern synthetic systems.

The future of synthetic respondents should not be about choosing between humans and AI. It should be about building transparent, scientifically grounded systems that combine the strengths of both while maintaining the standards of quality, representativeness, and rigor that social science research requires. This is why we are exploring how synthetic respondents might augment—rather than replace—probability-based panels in carefully defined ways. Potential applications for synthetic respondents include questionnaire testing, methodological experimentation, strategic integration with human samples, and research involving specific use cases where the benefits can be demonstrated empirically. 

As our work continues, one conclusion is already clear: trustworthy synthetic respondents require trustworthy human data. The better our human data infrastructure, the better our synthetic systems can become. And that means investing in high-quality, probability-based data collection remains as important as ever.


Suggested Citation

Christian, L. (2026, September 10). Human Data Remain Essential in the Age of Synthetic Respondents. [Web blog post]. NORC at the University of Chicago. Retrieved from www.norc.org.


Tags

Research Divisions

Departments, Centers & Programs



Solutions

Experts

Explore NORC Research Science Projects

Analyzing Parent Narratives to Create Parent Gauge™

Helping Head Start build a tool to assess parent, family, and community engagement

Client:

National Head Start Association, Ford Foundation, Rainin Foundation, Region V Head Start Association

America in One Room

A “deliberative polling” experiment to bridge American partisanship

Client:

Stanford University