首个系统评估大模型行为科学能力的基准,揭示其在群体层面表现的关键差距。
BehaviorBench: Benchmarking Foundation Models for Behavioral Science Tasks

- 构建四维评测体系,涵盖行为预测、决策、特质推断与知识应用
- 发现通用模型擅长个体预测,而专用模型更贴近真实人群分布
- 提出新模型Be.FM-1.5,在群体一致性上领先,适合行为科学研究
大模型在心理学、社会学、经济学等行为科学领域应用日益广泛。尽管其在问卷回答预测和人类实验模拟等任务中表现出色,但缺乏对跨任务、跨情境、跨人群表现的系统评估。本文提出BehaviorBench,一个全面的基准测试,从四大核心能力评估模型:(1)行为预测与模拟,(2)策略性决策,(3)受试者特质推断,(4)行为知识应用。关键在于,BehaviorBench同时评估个体与群体层面的表现,不仅关注单个受试者准确性,还衡量整体分布对齐度,这是行为有效性的重要要求。基于该基准,我们进一步开发了Be.FM-1.5,一个在行为数据上微调的专用模型。结果表明:商业通用模型在个体预测和知识密集型任务上占优,而行为基础模型在群体分布对齐上显著更优。值得注意的是,Be.FM-1.5在分布指标上领先,且在个体指标上保持竞争力,说明针对性行为适配可缩小差距。研究强调分布评估的重要性,确立BehaviorBench作为发展与评估行为对齐AI系统的基石,并展示Be.FM-1.5在广泛行为科学研究中的潜力。相关资源可通过 https://umich-foreseer.github.io/behaviorbench/ 获取。
原文摘要 · Abstract (English)
Foundation models have been increasingly applied to behavioral science domains such as psychology, sociology, and economics. While these models show promise in individual tasks such as survey response prediction and human-subject experiment simulation, there remains no systematic understanding of how well they perform across diverse behavioral science tasks, contexts, and populations. We introduce BehaviorBench, a comprehensive benchmark that evaluates foundation models along four core capabilities: (1) behavior prediction and simulation, (2) strategic decision-making, (3) subject-trait inference, and (4) behavioral knowledge application. Crucially, BehaviorBench evaluates model outputs at both the individual and distributional levels, capturing not only per-subject accuracy but also population-level alignment, an essential requirement for behavioral validity. Leveraging the tasks in BehaviorBench, we further develop Be.FM-1.5, extending the Be.FM family of behavioral foundation models fine-tuned on behavioral data. Our results reveal a considerable gap: proprietary general-purpose models excel at individual-level prediction and knowledge-intensive tasks, whereas behavioral foundation models, fine-tuned on behavioral data, achieve substantially stronger distributional alignment. Notably, Be.FM-1.5 leads on distributional metrics and remains competitive on individual-level metrics, suggesting that proper behavioral adaptation can close the gap. Our results highlight the importance of distributional evaluation, establish BehaviorBench as a foundation for developing and assessing behaviorally aligned AI systems, and demonstrate Be.FM-1.5's potential for a broad range of behavioral science studies. Our BehaviorBench and Be.FM-1.5 models can be accessed via https://umich-foreseer.github.io/behaviorbench/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。