arXiv:2605.23949cs.MAcs.AI2026-05

提出SODE框架,评估大模型在社会互动中的合作机制。

SODE: Analyzing Social Dynamics in LLM Agents

论文配图:SODE: Analyzing Social Dynamics in LLM Agents
图 1 · 摘自论文原文
  • 从直接互惠、间接互惠、群体动态三维度评估模型行为
  • 指令微调模型易被利用,推理模型倾向短期优化
  • 长周期框架可激发推理模型的互惠能力,适合对齐研究者

随着大语言模型演变为交互式智能体,理解其在人类社会动态中的行为对齐至关重要。尽管行为博弈论提供了研究框架,但以往工作多依赖平均分等结果指标,忽略了可持续合作背后的机制差异——相同得分可能源于截然不同的策略。为此,我们提出SODE(社会动态评估)框架,从三个演化维度评估LLM智能体:直接互惠(策略适应)、间接互惠(声誉敏感)与群体动态(合作韧性)。应用SODE发现系统性差异:指令微调模型常表现出‘被动顺从’,易遭剥削;而推理模型更关注短期优化,破坏长期合作。值得注意的是,采用‘长期视角框架’可激活推理模型的互惠能力。SODE为对齐人工智能智能体与复杂人类社会动态提供了一种机制驱动的系统性基准。

原文摘要 · Abstract (English)

As Large Language Models (LLMs) evolve into interactive agents, understanding their behavioral alignment within human social dynamics becomes essential. While behavioral game theory offers a framework to study these interactions, previous work has predominantly relied on outcome-based metrics such as average scores. This focus overlooks the mechanisms that facilitate sustainable cooperation, as identical scores can be derived from vastly different strategies. To bridge this gap, we introduce SODE (Social Dynamics Evaluation), a framework that evaluates LLM agents across three evolutionary dimensions: Direct Reciprocity for strategy adaptation, Indirect Reciprocity for reputation sensitivity, and Group Dynamics for cooperative resilience. Applying SODE reveals systematic divergences: instruction-tuned models often exhibit "passive compliance" that renders them vulnerable to exploitation, while reasoning models prioritize short-horizon optimization, destabilizing long-term cooperation. Notably, we demonstrate that a "long-horizon framing" can unlock reciprocal capabilities in reasoning models. Thus, SODE offers a systematic, mechanism-grounded benchmark for aligning AI agents with complex human social dynamics.

社会智能大模型对齐博弈论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。