arXiv:2609.04855cs.CLcs.AI2026-09

评测大模型跨文化调解能力,发现其介入时机与策略均有短板。

CC-Mediation: Evaluating Large Language Models for Cross-Cultural Conflict Mediation

论文配图:CC-Mediation: Evaluating Large Language Models for Cross-Cultural Conflict Mediation
图 1 · 摘自论文原文
  • 构建1661组跨文化对话数据集,基于敏感性发展模型设计
  • 提出两种新评估指标,与人类判断高度一致
  • 揭示大模型问题:介入过早或过晚,策略生成易失效

大语言模型在跨文化调解中需决定何时介入及如何回应。当前进展受限于缺乏具有可测量下游影响的调解数据集和严谨的跨文化立场变化评估指标。为此,我们提出CC-Mediation,一个包含1,661组十轮对话的跨文化调解基准,基于发展性跨文化敏感性模型(DMIS),涵盖文化冲突、调解干预及干预后演变轨迹。我们进一步提出两个基于DMIS的评估指标:轨迹AUC,衡量跨文化改善的持续性;带符号的Wasserstein-1距离,衡量文化立场转变的幅度与方向。两者均与人类对跨文化立场变化的判断高度一致。使用该基准发现,当前大模型在两方面存在局限:介入时机失败源于位置先验忽略对话内容,调解策略失败源于高层特征提取崩溃而非知识不足。

原文摘要 · Abstract (English)

Cross-cultural mediation by large language models (LLMs) requires deciding both when to intervene and how to respond in culturally grounded conflicts. Progress on this problem has been limited by the lack of (1) mediation datasets with measurable downstream effects and (2) principled metrics for evaluating intercultural stance change. To address these gaps, we introduce CC-Mediation, a cross-cultural mediation benchmark of $1{,}661$ ten-turn dialogues grounded in the Developmental Model of Intercultural Sensitivity (DMIS), containing culturally grounded conflicts, mediation interventions, and post-intervention trajectories. We further propose two DMIS-based evaluation metrics: Trajectory AUC, which measures the persistence of intercultural improvement over time, and a signed Wasserstein-1 distance, which measures the magnitude and direction of shifts in intercultural stance. Both metrics show strong agreement with human judgment of DMIS-grounded stance shift. Using CC-Mediation, we find that current LLMs have limitations on both axes: intervention timing (when) failure stems from a positional prior that ignores dialogue content, while mediation strategy (how) failure arises from a late-layer elicitation collapse rather than a knowledge deficit.

大模型跨文化调解评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。