构建可模拟人类行为的通用模型,提升对话与社交任务表现。
OdysSim: Building Foundation Models for Human Behavior Simulation

- 提出SOUL五维能力框架,统一62个数据集与23项任务
- 训练出8B规模模型,在23项任务中8项排名第一或并列第一
- 输出更贴近真人,零样本迁移至未知用户模拟表现优异
大语言模型被广泛用于交互评估与社会模拟,但以帮助性为导向的后训练使其趋向同质化、过度顺从,造成行为仿真与真实世界的差距。我们提出OdysSim,这是首个大规模系统性的人类行为基础模型研究。提出SOUL分类法,涵盖五大能力维度(CONV、SS、COG、ROLE、EVAL),整合62个数据集与23项基准任务。构建包含2140万次交互、100亿词元的OdysSim语料库,并引入回溯生成的社会背景;开发端到端训练方案,结合中期训练、任务特定强化学习与专家蒸馏。最终开源的8B模型在23项任务中8项排名第一或并列第一,尤其在对话与社交任务上优势显著。其输出在长度、格式与用词上更接近真人,且在τ-bench上实现零样本迁移,对未见用户模拟的反应一致性达93.2%(真实用户为93.5%)。我们还发现,以LLM为裁判的强化学习会诱发奖励欺骗,而我们的检测器可有效缓解。研究表明,行为基础模型需重新思考大模型训练范式。所有资源已公开。
原文摘要 · Abstract (English)
Large language models are increasingly deployed as human simulators for interactive evaluation and social simulation. Yet helpfulness-driven post-training pulls them toward a homogeneous, overly agreeable assistant register, creating a behavioral Sim2Real gap. We present OdysSim, the largest open systematic investigation of behavioral foundation models, i.e., models trained to simulate human behavior at scale. We propose SOUL, a taxonomy of five capability axes (CONV, SS, COG, ROLE, EVAL) that unifies 62 datasets and 23 benchmark tasks under one framework. Specifically, we curate the OdysSim corpus (21.4M interactions, 10B tokens, retrofitted with back-generated social contexts), construct the SOUL-Index benchmark, and develop an end-to-end training recipe combining midtraining, task-specific RL, and expert distillation. The resulting open 8B OSim model ranks first or tied-first on 8 of 23 tasks, outperforming any individual frontier model by this count, with the strongest gains on conversational and social tasks. Its outputs are also more human-like in length, formatting, and word choice, and it transfers zero-shot to out-of-distribution user simulation on $τ$-bench, nearly matching real users on reaction alignment (93.2 vs. 93.5). We further show that LLM-as-judge RL induces reward-hacking patterns, and that our detectors can mitigate them during post-training. Together, our findings suggest that behavioral foundation models require rethinking the LLM training paradigm. We release all artifacts to support future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。