arXiv:2606.05563cs.AIcs.CL2026-06

构建跨领域真实调解场景,评估大模型在复杂社会认知下的主动调解能力。

SoCRATES: Towards Reliable Automated Evaluation of Proactive LLM Mediation across Domains and Socio-cognitive Variations

论文配图:SoCRATES: Towards Reliable Automated Evaluation of Proactive LLM Mediation across Domains and Socio-cognitive Variations
图 1 · 摘自论文原文
  • 基于真实冲突构建八领域动态场景,模拟多方情绪与意图变化。
  • 引入局部化评估器,对推进话题的对话回合精准打分,准确率达0.82(人类专家对齐)。
  • 揭示当前最强模型仅能缩小三分之一未调解共识差距,社会适应是关键瓶颈。

评估大语言模型(LLM)调解者仍具挑战性,因调解过程受争议方情绪、意图与上下文动态影响,呈实时演化轨迹。现有测试集仅依赖少数专家编写的领域,主要变化策略姿态,且对每轮对话在所有话题上评分,引入大量无关话题噪声。我们提出SoCRATES,一个面向现实多领域情境的主动式LLM调解评估基准。该基准通过智能体流水线从真实冲突中构建八领域场景,涵盖五类社会认知适应维度(策略姿态、参与方构成、历史长度、情绪反应性、文化身份),并采用话题局部化评估器,仅对推进特定话题的对话回合进行评分。评估器与人类专家达成0.82一致性,超过逐轮基线两倍以上。对八个前沿大模型的基准测试显示,即使最强模型在多样且真实的测试环境下,也仅能缩小约三分之一未调解共识差距,且性能随社会认知维度显著波动,凸显社会适应能力是未来进展的核心。

原文摘要 · Abstract (English)

Evaluating LLM mediators remains challenging, as mediation unfolds as a real-time trajectory shaped by disputants' shifting emotions, intentions, and context. Existing testbeds rely on a few expert-authored domains, vary mainly strategic posture, and score every turn against every topic, introducing off-topic noise. We introduce SoCRATES, a benchmark for evaluating proactive LLM mediators in realistic, multi-domain testbeds. It constructs scenarios from real conflicts through an agentic pipeline across eight domains, probes five socio-cognitive adaptation axes (strategic posture, party composition, history length, emotional reactivity, and cultural identity), and scores each topic only on the turns that advance it via a topic-localized evaluator. The evaluator reaches 0.82 alignment with human experts, more than doubling a per-turn baseline. Benchmarking eight frontier LLMs, we find that even the strongest mediator closes only about a third of the unmediated consensus gap under diverse and realistic testbeds, with performance varying sharply by socio-cognitive axis, highlighting that progress lies in social adaptation to diverse conditions.

大模型评估主动调解社会认知多领域

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。