arXiv:2502.14119cs.CL2025-02ACL被引 1

用指代可及性测试评估模型对篇章理解的能力,发现大模型与人类在理解上存在差异。

Meaning Beyond Truth Conditions: Evaluating Discourse Level Understanding via Anaphora Accessibility

  • 提出指代可及性任务,诊断模型对篇章层面的理解能力
  • 大模型与人类在部分任务上表现一致,但在结构抽象理解上差距明显
  • 适合关注语言认知机制和模型推理局限的研究者阅读

我们提出一个自然语言理解能力的层级框架,主张应从词法和句子层面推进到篇章层面。为此,我们设计了指代可及性任务作为篇章理解的诊断工具,并构建了一个受动态语义理论启发的评估数据集。我们评估了人类与大语言模型在此数据集上的表现,发现两者在某些任务上一致,但在其他任务上出现分歧。这种分歧可归因于大语言模型在语言理解中更依赖特定词汇项,而人类则更敏感于结构抽象。该研究为理解大模型的语言认知机制提供了新视角。

原文摘要 · Abstract (English)

We present a hierarchy of natural language understanding abilities and argue for the importance of moving beyond assessments of understanding at the lexical and sentence levels to the discourse level. We propose the task of anaphora accessibility as a diagnostic for assessing discourse understanding, and to this end, present an evaluation dataset inspired by theoretical research in dynamic semantics. We evaluate human and LLM performance on our dataset and find that LLMs and humans align on some tasks and diverge on others. Such divergence can be explained by LLMs' reliance on specific lexical items during language comprehension, in contrast to human sensitivity to structural abstractions.

篇章理解指代消解语言模型认知机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。