大模型在指代消解中难以同时准确判断指代关系与识别歧义。
Correct-Detect: Balancing Performance and Ambiguity Through the Lens of Coreference Resolution in LLMs
- 发现大模型在指代消解和歧义检测间存在内在权衡
- 仅用少量提示时,模型可达成较好性能但无法兼顾两者
- 揭示了语言理解中隐式能力的潜在矛盾,适合研究模型认知机制者阅读
大型语言模型(LLMs)旨在体现人类的语言能力。然而,人类具备广泛且具身化的上下文知识,这对检测和解决语言歧义至关重要,即使在孤立文本片段中也是如此。语义歧义的一个基础案例是核心指代消解任务:代词与之前提及的人称如何关联?这一能力几乎隐含于每个下游任务中,该层级的歧义存在会显著影响模型表现。我们发现,尽管模型在无须额外提示的情况下,能在核心指代消解和歧义检测任务中均取得良好表现,但无法同时兼顾两者。本文提出 CORRECT-DETECT 二元权衡:虽然模型具备这两种能力并隐式运用,但成功平衡二者仍极具挑战。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are intended to reflect human linguistic competencies. But humans have access to a broad and embodied context, which is key in detecting and resolving linguistic ambiguities, even in isolated text spans. A foundational case of semantic ambiguity is found in the task of coreference resolution: how is a pronoun related to an earlier person mention? This capability is implicit in nearly every downstream task, and the presence of ambiguity at this level can alter performance significantly. We show that LLMs can achieve good performance with minimal prompting in both coreference disambiguation and the detection of ambiguity in coreference, however, they cannot do both at the same time. We present the CORRECT-DETECT trade-off: though models have both capabilities and deploy them implicitly, successful performance balancing these two abilities remains elusive.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。