arXiv:2506.01089cs.CLcs.AI2025-06ACL被引 1

测试大模型对'I'、'你'等指代词的理解能力,发现表现差异大。

Un-considering Contextual Information: Assessing LLMs' Understanding of Indexical Elements

  • 构建首个英文指代词数据集,含1600道多选题
  • GPT-4o等模型对'I'理解较好,对'you'等表现差
  • 引号等语法线索有时助益,有时反而干扰

大型语言模型在指代消解任务中表现优异,但以往研究主要聚焦于名词和第三人称代词。本研究首次评估大模型对英语中具有独特语言特性的指代词(如I、you、here、tomorrow)的理解能力,发布包含1600个多项选择题的英文指代词数据集。评估了GPT-4o、Claude 3.5 Sonnet、Gemini 1.5 Pro和DeepSeek V3等前沿模型。结果表明,模型对'I'有较好表现,但对'you'、'here'、'tomorrow'等则存在困难;句法线索(如引号)在某些情况下提升性能,但在其他情况下反而降低表现。代码与数据已开源。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated impressive performances in tasks related to coreference resolution. However, previous studies mostly assessed LLM performance on coreference resolution with nouns and third person pronouns. This study evaluates LLM performance on coreference resolution with indexical like I, you, here and tomorrow, which come with unique challenges due to their linguistic properties. We present the first study examining how LLMs interpret indexicals in English, releasing the English Indexical Dataset with 1600 multiple-choice questions. We evaluate pioneering LLMs, including GPT-4o, Claude 3.5 Sonnet, Gemini 1.5 Pro, and DeepSeek V3. Our results reveal that LLMs exhibit an impressive performance with some indexicals (I), while struggling with others (you, here, tomorrow), and that syntactic cues (e.g. quotation) contribute to LLM performance with some indexicals, while they reduce performance with others. Code and data are available at: https://github.com/metehanoguzz/LLMs-Indexicals-English.

大模型评测指代消解自然语言理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。