AutoPersonas: The Bottleneck of Agent Simulation Shifts from Scale to Role Diversity and Repetition Contamination
The AutoPersonas preprint, submitted on July 9, 2026, analyzes 1,600 personas generated by 8 models and discusses issues of repetition, bias, and effective diversity in automatic persona generation.
Using large numbers of virtual users or expert agents for product research may seem to offer low-cost scaling, but the number of personas does not equal coverage of viewpoints. This study turns persona generation itself into a measurable object, revealing that repetitive templates and model biases can create a false 'user consensus.'
The paper compares semantic coverage, repetition patterns, and demographic distribution of personas generated by multiple models. The core method transforms the quality of persona sets from subjective perception into diversity and redundancy metrics. For simulation systems, sampling strategies, deduplication, and calibration with real users are as important as downstream agent capabilities.
Synthetic user research, market simulation, and multi-agent social experiments require new data quality standards; persona pools without coverage and bias audits cannot replace real interviews or be packaged as market evidence.
When using synthetic personas for early exploration, retain a real sample calibration set, disclose the model and sampling methods, and incorporate deduplication rates, coverage gaps, and deviations from real users into decision confidence.
The preprint's conclusions still need to be replicated across more languages, regions, and vertical populations, and to test whether persona diversity metrics can predict downstream decision quality, avoiding metrics that only reflect surface-level textual differences.