研究大模型如何理解复数指代歧义,发现其识别能力有限且不一致。
Can LLMs Detect Ambiguous Plural Reference? An Analysis of Split-Antecedent and Mereological Reference
- 通过提示词和预测任务测试大模型对复数指代的处理机制
- 模型能察觉部分歧义指代,但无法稳定识别所有歧义情况
- 适合关注语言理解偏差与人类认知差异的研究者阅读
本研究旨在探究大语言模型(LLMs)在模糊与明确语境中对复数指代的表征与理解能力。核心问题包括:(1) 大模型是否表现出类似人类的复数指代偏好?(2) 大模型能否检测复数回指表达中的歧义并识别可能指代对象?为此,我们设计了一系列实验,涵盖基于下一个词预测的任务中的代词生成、代词解释以及不同提示策略下的歧义检测。结果表明,大模型有时能意识到模糊代词的潜在指代对象,但在选择解释时并不总遵循人类习惯,尤其当可能指代未被明确提及。此外,模型在缺乏直接指令时难以识别歧义。研究还发现不同实验类型间结果存在不一致性。
原文摘要 · Abstract (English)
Our goal is to study how LLMs represent and interpret plural reference in ambiguous and unambiguous contexts. We ask the following research questions: (1) Do LLMs exhibit human-like preferences in representing plural reference? (2) Are LLMs able to detect ambiguity in plural anaphoric expressions and identify possible referents? To address these questions, we design a set of experiments, examining pronoun production using next-token prediction tasks, pronoun interpretation, and ambiguity detection using different prompting strategies. We then assess how comparable LLMs are to humans in formulating and interpreting plural reference. We find that LLMs are sometimes aware of possible referents of ambiguous pronouns. However, they do not always follow human reference when choosing between interpretations, especially when the possible interpretation is not explicitly mentioned. In addition, they struggle to identify ambiguity without direct instruction. Our findings also reveal inconsistencies in the results across different types of experiments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。