arXiv:2602.10635cs.AIcs.LG2026-02中稿 · ICML被引 1

让AI更懂复杂社交行为,提升跨场景泛化能力。

OmniSapiens: A Foundation Model for Social Behavior Processing via Heterogeneity-Aware Relative Policy Optimization

  • 用新强化学习方法平衡异构数据的学习信号。
  • 在10个行为任务上表现最佳,最高提升12.02%。
  • 推理过程更稳定可解释,适合真实社交应用。

社会智能人工智能系统需在多样人类行为任务中推理并泛化至新社交情境。然而,行为数据本质异构,包含多种模态与预测目标,导致样本间训练信号不均,引发学习动态失衡,挑战现有模型。为此,我们提出用于社会行为处理的奠基模型 Omnisapiens-7B 2.0,通过异构感知相对策略优化(Heterogeneity-Aware Relative Policy Optimization)显式应对异构行为数据学习问题。该方法通过近似每个样本对策略更新的贡献,驱动几何中心化、惯性平滑的优势调节,实现稳定训练。Omnisapiens-7B 2.0 在10个行为任务上取得最优且最一致的表现,同时在全部5个保留基准测试中均达到最佳成绩,性能提升最高达+12.02%和+9.37%。此外,其推理轨迹更一致且可解释,支持可靠的实际行为应用。模型代码已公开于 https://github.com/MIT-MI/human_behavior_atlas。

原文摘要 · Abstract (English)

Socially intelligent AI systems must reason across diverse human behavioral tasks and generalize to new social contexts. However, behavioral data is inherently heterogeneous, comprising diverse modalities and prediction targets that produce uneven training signals across samples, creating imbalanced learning dynamics that challenge existing AI models. To address this, we develop Omnisapiens-7B 2.0, a foundation model for social behavior processing that explicitly addresses learning from heterogeneous behavioral data. This is enabled through Heterogeneity-Aware Relative Policy Optimization, a new RL method that rebalances learning signals across samples by approximating each sample's contribution to the policy update and using these estimates to drive geometrically centered, inertially smoothed advantage modulation for stable training. Omnisapiens-7B 2.0 achieves the best and most consistent performance across 10 behavioral tasks, while also attaining the best performance on all five held-out benchmarks, with gains of up to +12.02% and +9.37% respectively. Furthermore, it demonstrates more consistent and interpretable reasoning traces, supporting reliable real-world behavioral applications. Our model is available at https://github.com/MIT-MI/human_behavior_atlas.

社会行为强化学习基础模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。