构建统一行为基准,提升模型对心理社会行为的理解能力
Human Behavior Atlas: Benchmarking Unified Psychological and Social Behavior Understanding
- 整合多模态行为数据,构建跨任务统一基准
- 在10万+样本上训练的模型显著优于现有多模态大模型
- 适合心理学、人机交互与通用行为理解研究者使用
利用智能系统感知心理与社会行为——即通过可观察行为和社交互动展现的情感、认知及病理状态——仍面临挑战,因其复杂、多维且高度个性化。现有工作依赖专用数据集与单任务系统,难以实现可扩展性、跨任务迁移与广泛泛化。为此,我们构建了人类行为图谱(Human Behavior Atlas),一个统一的多样化行为任务基准,旨在支持心理与社会行为理解基础模型的发展。该基准包含超过10万条文本、音频与视觉模态样本,涵盖情感状态、认知状态、病理特征与社会过程等任务。统一设计可减少冗余与成本,支持跨任务高效训练,并增强行为特征在不同领域的泛化能力。我们在该基准上训练了三个模型:Omnisapiens-7B SFT、Omnisapiens-7B BAM 和 Omnisapiens-7B RL。结果表明,基于该基准训练的模型在多种行为任务中持续优于现有多模态大模型。此外,预训练于该基准还提升了模型在新行为数据集上的迁移性能;通过针对性使用行为描述符,可获得显著性能提升。基准、模型与代码详见:https://github.com/MIT-MI/human_behavior_atlas。
原文摘要 · Abstract (English)
Using intelligent systems to perceive psychological and social behaviors, that is, the underlying affective, cognitive, and pathological states that are manifested through observable behaviors and social interactions, remains a challenge due to their complex, multifaceted, and personalized nature. Existing work tackling these dimensions through specialized datasets and single-task systems often miss opportunities for scalability, cross-task transfer, and broader generalization. To address this gap, we curate Human Behavior Atlas, a unified benchmark of diverse behavioral tasks designed to support the development of foundation models for understanding psychological and social behaviors. Human Behavior Atlas comprises over 100,000 samples spanning text, audio, and visual modalities, covering tasks on affective states, cognitive states, pathologies, and social processes. Our unification efforts can reduce redundancy and cost, enable training to scale efficiently across tasks, and enhance generalization of behavioral features across domains. On Human Behavior Atlas, we train three models: Omnisapiens-7B SFT, Omnisapiens-7B BAM, and Omnisapiens-7B RL. We show that training on Human Behavior Atlas enables models to consistently outperform existing multimodal LLMs across diverse behavioral tasks. Pretraining on Human Behavior Atlas also improves transfer to novel behavioral datasets; with the targeted use of behavioral descriptors yielding meaningful performance gains. The benchmark, models, and codes can be found at: https://github.com/MIT-MI/human_behavior_atlas.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。