用83亿虚拟人模拟真实用户行为,高效评估AI产品体验。
MatrAIx: Simulating the World with 8.3 Billion Persona Agents

- 构建8.3亿带属性的虚拟人格数据库,支持多样用户模拟。
- 在4类环境中完成18,189次测试,验证不同背景用户的行为差异。
- 适合需要大规模用户反馈的AI产品、数字服务开发者使用。
当前AI系统和数字产品的用户体验评估成本高、难扩展。为解决此问题,本文提出MatrAIx——一个基于大规模模拟用户群体的评估框架。其核心包括:第一,Persona 8B包含8.3亿个由1,290个类别维度描述的人格记录,其中约100万条经质量筛选,包含599,847条真人来源与40万条合成数据;第二,提供调查、聊天机器人、网页和应用四个交互环境;第三,涵盖1,010项跨25个领域的任务。实验共开展18,189次评估,使用Claude Opus 4.8、GPT 5.5和Claude Haiku 4.5三款大模型驱动代理。结果揭示了用户对价格变动的犹豫、对失败助手的容忍度及响应延迟偏好等差异。两项验证研究显示:在400次受控测试中,91.5%的行为表现符合预设人格特征;人类与大模型评委均确认真人来源人格的提取质量可靠。该框架实现了从人格生成到评估反馈的端到端闭环。
原文摘要 · Abstract (English)
Human evaluation of AI systems and digital products is costly, slow, and difficult to scale. Offline evaluations are more scalable but often abstract away human diversity and interactive behavior. We therefore introduce MatrAIx, a population-scale simulated-user evaluation infrastructure for testing AI systems and digital products with heterogeneous users. MatrAIx has three core components: First, Persona 8B contains 8.3 billion persona records represented by 1,290 categorical dimensions. Records are either sampled from a dependency graph that preserves correlated attributes or derived from human-authored profiles. We release a quality-filtered coreset of approximately 1 million personas, comprising 599,847 human-grounded and 400,000 synthetic records. Second, the MatrAIx Playground provides four environments in which diverse users evaluate and interact with digital products: Survey, AI Chatbot, Web, and App. Third, MatrAIx provides 1,010 application tasks spanning more than 25 domains, including Commerce, Software, Finance, and Healthcare. We conducted 18,189 evaluation trials across eight representative tasks. Persona agents were powered by three LLMs: Claude Opus 4.8, GPT 5.5, and Claude Haiku 4.5. The resulting feedback captures how decisions and preferences vary across persona backgrounds, including hesitation after a price increase, willingness to continue after an AI assistant fails, and latency tolerance. We conducted two main validation studies: First, a 400-trial controlled study evaluated persona adherence across ten behavioral attributes and all four environments. The declared behavior was expressed or correctly suppressed in 366 trials (91.5%). Second, human and LLM judges evaluated the extraction quality of human-grounded personas. Overall, MatrAIx provides an end-to-end infrastructure for evaluating AI systems and digital products with diverse simulated human users.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。