arXiv:2607.23740cs.CL2026-07

给大模型装上社会智能,让它们懂人心、守规矩、会变通。

ZenGen: Social Mind for LLMs

论文配图:ZenGen: Social Mind for LLMs
图 1 · 摘自论文原文
  • 构建心理基础的评测框架SoMBench,覆盖71种社交任务场景。
  • 最优模型仅达72.08%准确率,17个次级维度均未接近90%天花板。
  • 通过四类运行时支持,显著提升大模型的社会适应能力。

随着大语言模型从孤立任务转向长期人机环境服务,需具备社会智能——即推断心理状态、追踪社会关系、推理规范并根据上下文调整行为的能力。本文提出ZenGen框架,涵盖社会智能的测评、内化与部署阶段。测评方面,引入基于心理学的SoMBench基准,覆盖3个主维度、17个次级维度和71种任务范式,包含284个共享场景与3,481个专家验证实例,控制提问格式、叙事视角与上下文长度。对20个代表性模型评估显示,最优模型整体准确率仅为72.08%,且无一二级维度进入90%近天花板区间。内化方面,开发诊断驱动的训练方案ZenGen,融合监督微调、在线策略蒸馏与基于评分标准的强化学习,在五个社交认知基准上持续优于基线模型,其中ZenGen-27B-Stage2平均表现最佳,ZenGen-32B-Stage2与DeepSeek-V4-Pro相当。部署阶段,构建Actio推理架构,集成四类运行时支持:PRISM(程序引导)、Starling(实时心理表征)、SAGE(可复用经验)与门控RAG(外部社会与规范知识)。在五种基座模型与三个基准上,全组合架设使14/15模型-基准对表现提升,8组达到最优或并列最优,验证了类型化运行时支持的有效性。结果表明,社会智能大模型需在评估、参数内化与运行时接地三方面协同推进。

原文摘要 · Abstract (English)

As large language models move from isolated task solving toward long-term service in human environments, they require social intelligence: the ability to infer mental states, track social relations, reason over norms, and adapt behavior under context. This report presents ZenGen, an integrated framework for measuring, internalizing, and grounding social intelligence. For measurement, we introduce SoMBench, a psychology-grounded benchmark spanning 3 primary dimensions, 17 secondary dimensions, and 71 task paradigms. It controls question format, narrative perspective, and context length across 284 shared scenarios and 3,481 expert-verified instances. Evaluation of 20 representative LLMs reveals substantial headroom: the best model achieves only 72.08% overall accuracy, and none of the 17 secondary dimensions reaches the 90% near-ceiling band. For internalization, we develop ZenGen, a diagnosis-driven training recipe combining supervised fine-tuning, on-policy distillation, and rubric-based reinforcement learning. Across five social-cognition benchmarks, ZenGen consistently outperforms its base models, with ZenGen-27B-Stage2 achieving the best average score and ZenGen-32B-Stage2 remaining competitive with DeepSeek-V4-Pro. For deployment-time grounding, we build Actio, a harness-controlled inference architecture that routes four typed supports into reasoning: PRISM for procedural guidance, Starling for runtime mental-state representation, SAGE for reusable experience, and gated RAG for external social and normative knowledge. Across five base models and three benchmarks, the full harness improves 14 of 15 model-benchmark pairs and is best or tied for best in 8, demonstrating the effectiveness of typed runtime support. Together, these results show that socially intelligent LLMs require coordinated advances in evaluation, parametric internalization, and deployment-time grounding.

社会智能评测基准运行时支持大模型训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。