arXiv:2603.09993cs.CLcs.AI2026-03

评测大模型在真实语境中的隐含意图理解能力

CEI: A Benchmark for Evaluating Pragmatic Reasoning in Language Models

  • 构建300个真人验证的语境场景,测试模型对隐含意义的推断
  • 涵盖讽刺、委婉、被动攻击等5类复杂语用现象,含权力关系设定
  • 标注者一致性低但合理,反映语用理解本就存在多元解读

语用推理——超越字面意义理解说话人真实意图——是日常交流的核心,却仍难为大语言模型所掌握。本文提出情境情感推理(CEI)基准:包含300个经人工验证的情境,用于评估大模型在复杂语用表达中的消歧能力。每个场景均包含情境背景、说话人与听者角色(明确标注权力关系),以及一句模糊表达。数据覆盖职场、家庭、社交和服务业中的五种语用类型(讽刺/反语、矛盾信号、策略性礼貌、被动攻击、转移话题/误导),并设置三种权力结构(平级、上级对下级、下级对上级)。三名训练过的标注员独立标注每条数据。标注者间一致性(Fleiss' kappa = 0.06-0.25,按子类型)较低,但属预期结果:语用推理允许多种合理解释,这种分歧本身具有信息价值。本文详述标注方法,包括结合自动化统计检测与专家仲裁的四级质量控制流程。CEI已开放获取,采用CC-BY-4.0许可。

原文摘要 · Abstract (English)

Pragmatic reasoning, inferring intended meaning beyond literal semantics, underpins everyday communication yet remains difficult for large language models. We present the Contextual Emotional Inference (CEI) Benchmark: 300 human-validated scenarios for evaluating how well LLMs disambiguate pragmatically complex utterances. Each scenario pairs a situational context and speaker-listener roles (with explicit power relations) against an ambiguous utterance. The dataset covers five pragmatic subtypes (sarcasm/irony, mixed signals, strategic politeness, passive aggression, deflection/misdirection) drawn from workplace, family, social, and service settings, with three power configurations (peer, higher-to-lower, lower-to-higher). Three trained annotators independently labeled every scenario. Inter-annotator agreement (Fleiss' kappa = 0.06-0.25 by subtype) is low but expected: pragmatic inference admits multiple valid readings, and the disagreement itself is informative. We describe our annotation methodology, including a 4-level quality control pipeline that combines automated statistical checks with expert adjudication. CEI is released under CC-BY-4.0.

语用推理大模型评测人类标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。