arXiv:2606.21315cs.AI2026-06

构建可持续学习的社会智能模型,让小模型长期提升社交协作能力。

Social World Model for Lifelong Social Intelligence

论文配图:Social World Model for Lifelong Social Intelligence
图 1 · 摘自论文原文
  • 将社交互动分解为五维闭环框架,生成可迭代的学习信号。
  • 在ASCENT-Bench上,模型完成率媲美闭源Gemini,零遗忘且通过率更高。
  • 适合追求长期进化能力的开源模型研究者与智能体开发者。

社交智能是语言智能体的核心能力,但现有研究多聚焦静态评估,忽视其持续发展与积累。当前存在两大问题:社交交互轨迹缺乏统一结构化表示,难以形成可迭代的学习信号;能力提升与保留常被孤立研究,阻碍对持续演化的评估。为此,我们提出社交世界模型(Social World Model),将社交互动分解为场景设置、观察、心理状态、行动和对话五个维度,构建闭环学习框架。智能体收集交互经验,转化为偏好信号用于模型更新,并部署新策略继续学习。此外,提供可复用的数据合成机制与终身学习基准,将社交能力从‘评估对象’转变为‘可持续训练对象’。在ASCENT-Bench上的验证表明,交互式训练的Qwen2.5-7B模型在全部五项核心指标上优于基线,完成率媲美闭源Gemini 3 Flash,通过率更高,且在三个难度等级下实现零遗忘。该端到端方法提供可训练、可验证、可保留的路径,证明小型开源模型也能持续获得具备竞争力的社交协调能力。

原文摘要 · Abstract (English)

Social intelligence is a core competency for language agents, yet current research primarily focuses on static capability evaluation rather than how these skills are continuously shaped and accumulated. This gap calls for a shift toward sustainable learning paradigms. Currently, two methodological pain points exist: social interaction trajectories lack unified structured representations to form iterable learning signals, and capability improvement and retention are typically studied in isolation, hindering the assessment of continuous evolution. To bridge this gap, we propose the Social World Model. We decompose social interaction into five dimensions (scene setting, observation, mental state, action, and dialogue) to build a closed-loop learning framework. In this setup, agents collect interaction experiences, convert them into preference signals for model updating, and redeploy the updated policy for continued learning. Additionally, we provide a reusable data synthesis mechanism and a lifelong learning benchmark, transforming social capabilities from an "object of evaluation" into an "object of sustainable training". Validating our framework on the ASCENT-Bench, the interactively trained Qwen2.5-7B model outperforms its baseline across all five core metrics. Notably, it matches the closed-source Gemini 3 Flash in completion rate, exceeds it in pass rate, and achieves zero forgetting across three difficulty levels. Unlike prior works that merely report static comparisons or capability decay, this end-to-end approach provides a trainable, verifiable, and retainable pathway, demonstrating that small open-source models can sustainably acquire competitive social coordination capabilities.

社交智能持续学习语言模型闭环训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。