arXiv:2602.05115cs.AIcs.CL2026-02被引 2

测试大模型在沟通障碍下的社交智能,发现性能普遍下降超45%。

SocialVeil: Probing Social Intelligence of Language Agents under Communication Barriers

  • 构建模拟认知差异导致沟通障碍的社交环境
  • 平均互理解降低45%,困惑度上升近50%
  • 适合研究大模型社交能力与适应策略的学者

大型语言模型(LLMs)在交互环境中评估其社交智能日益重要。然而,现有基准通常假设理想化通信,限制了我们诊断模型在更真实、不完美情境中维持和修复互动的能力。为弥补这一差距,我们提出 extsc{SocialVeil},一个可模拟因认知差异引发沟通障碍的社交学习环境。基于对人类互动中沟通挑战的系统文献综述, extsc{SocialVeil} 引入三种典型干扰:语义模糊、社会文化错配与情感干扰。我们还引入两个障碍感知评价指标——未解决困惑与相互理解,用于评估受损通信下的互动质量。在720个场景中对四款前沿LLM的实验表明,障碍显著损害性能,相互理解平均下降超过45%,困惑度提升近50%。人工评估验证了模拟障碍的真实性(ICC≈0.78,Pearson r≈0.80)。我们进一步证明,修复指令与交互式学习等适应策略效果有限,难以接近无障环境表现。本工作推动社交交互环境向现实世界通信更靠近,为探索LLM代理的社交智能开辟新路径。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly evaluated in interactive environments to test their social intelligence. However, existing benchmarks often assume idealized communication between agents, limiting our ability to diagnose whether LLMs can maintain and repair interactions in more realistic, imperfect settings. To close this gap, we present \textsc{SocialVeil}, a social learning environment that can simulate social interaction under cognitive-difference-induced communication barriers. Grounded in a systematic literature review of communication challenges in human interaction, \textsc{SocialVeil} introduces three representative types of such disruption, \emph{semantic vagueness}, \emph{sociocultural mismatch}, and \emph{emotional interference}. We also introduce two barrier-aware evaluation metrics, \emph{unresolved confusion} and \emph{mutual understanding}, to evaluate interaction quality under impaired communication. Experiments across 720 scenarios and four frontier LLMs show that barriers consistently impair performance, with mutual understanding reduced by over 45\% on average, and confusion elevated by nearly 50\%. Human evaluations validate the fidelity of these simulated barriers (ICC$\approx$0.78, Pearson r$\approx$0.80). We further demonstrate that adaptation strategies (Repair Instruction and Interactive learning) only have a modest effect far from barrier-free performance. This work takes a step toward bringing social interaction environments closer to real-world communication, opening opportunities for exploring the social intelligence of LLM agents.

社交智能沟通障碍大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。