用进化算法生成多样虚拟人物,覆盖罕见特质组合。
Persona Generators: Generating Diverse Synthetic Personas for Arbitrary Contexts
- 基于AlphaEvolve迭代优化,让LLM自动生成多样化虚拟角色
- 在6个维度上超越基线,显著提升长尾特质覆盖率
- 适合评估未来场景下AI对多元用户群体的响应能力
评估与人类交互的AI系统需要理解其在多样化人群中的表现,但获取代表性人类数据往往成本高昂或不可行,尤其针对新技术或未来假设场景。已有研究利用大语言模型生成高保真度的合成人格,准确复现特定个体信念与行为,但多数方法依赖详细目标人群数据,且侧重密度匹配(复制常见行为)而非支持覆盖(涵盖可能性),导致长尾行为未被充分探索。本文提出Persona Generators,可针对任意上下文生成多样化合成人群的函数。通过基于AlphaEvolve的迭代改进循环,以大语言模型作为突变算子,在数百次迭代中优化生成器代码。该过程产出轻量级生成器,能将少量描述自动扩展为覆盖关键多样性轴线上广泛意见与偏好的合成人群。实验表明,演化后的生成器在六项多样性指标上显著优于现有基线,生成了标准LLM输出难以实现的罕见特质组合。
原文摘要 · Abstract (English)
Evaluating AI systems that interact with humans requires understanding their behavior across diverse user populations, but collecting representative human data is often expensive or infeasible, particularly for novel technologies or hypothetical future scenarios. Recent work in Generative Agent-Based Modeling has shown that large language models can simulate human-like synthetic personas with high fidelity, accurately reproducing the beliefs and behaviors of specific individuals. However, most approaches require detailed data about target populations and often prioritize density matching (replicating what is most probable) rather than support coverage (spanning what is possible), leaving long-tail behaviors underexplored. We introduce Persona Generators, functions that can produce diverse synthetic populations tailored to arbitrary contexts. We apply an iterative improvement loop based on AlphaEvolve, using large language models as mutation operators to refine our Persona Generator code over hundreds of iterations. The optimization process produces lightweight Persona Generators that can automatically expand small descriptions into populations of diverse synthetic personas that maximize coverage of opinions and preferences along relevant diversity axes. We demonstrate that evolved generators substantially outperform existing baselines across six diversity metrics on held-out contexts, producing populations that span rare trait combinations difficult to achieve in standard LLM outputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。