arXiv:2607.25094cs.CLcs.AI2026-07

测试大模型理解言外之意更新的能力,发现其远未达到人类水平。

Evaluating Communicative Belief Updates in Large Language Models via Implicature Recognition and Cancellation

论文配图:Evaluating Communicative Belief Updates in Large Language Models via Implicature Recognition and Cancellation
图 1 · 摘自论文原文
  • 构建首个专家标注的言外之意取消数据集,用于评估模型对隐含信念的理解。
  • 模型在自然场景下的信念更新能力显著弱于人类,尤其在复杂语境中。
  • 模型依赖先验信念,且不同类型的信念更新失败模式各异,适合认知研究者参考。

人类语言依赖于未言明的信念及其更新,这对大语言模型(LLMs)与用户间的有效沟通至关重要。本文通过言外之意识别与取消任务,评估LLMs理解隐含信念及其更新的能力:即话语含义被削弱或否定的语用现象。我们创建了首个专家标注的言外之意取消数据集ImplicatureX,通过众包获取人类对言外之意及其取消判断的标注。结果表明,当前LLM在信念更新理解上仍落后于人类,尤其在更自然的语境中表现不佳。控制实验显示,模型的成功可能部分源于对先验信念的依赖,而失败则取决于信念类型和表达形式。整体而言,现有LLM尚未达到人类对未言明信念及更新的理解水平。代码与数据已公开于https://github.com/cesare-spinoso/ImplicatureX。

原文摘要 · Abstract (English)

Human language is driven by unspoken beliefs and belief updates, making these critical to model for successful communication between large language models (LLMs) and their users. In this paper, we evaluate the ability of LLMs to recognize unspoken beliefs made through implicatures and to understand their updates through implicature cancellation: the pragmatic phenomenon whereby an utterance's implied meaning is weakened or negated. We create the first expert-annotated implicature cancellation dataset, ImplicatureX, crowdsourced for human judgements of implicatures and their corresponding cancellations. We find that LLM belief update understanding lags behind that of humans, especially in more naturally-occurring scenarios. Additional control experiments suggest that successes in LLM belief updates may stem in part from a reliance on prior beliefs, and that failures in belief updates may depend on their type and on their form. Overall, our study suggests that current LLMs have not yet reached human-level understanding of unspoken beliefs and belief updates. Code and data are available at https://github.com/cesare-spinoso/ImplicatureX.

语言理解信念更新言外之意大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。