arXiv:2502.05793cs.CL2025-02NAACL被引 1

发现现有NLI模型在跨上下文推理时严重失效,提出新基准测试其表现。

On Reference (In-)Determinacy in Natural Language Inference

  • 构建跨上下文匹配的诊断数据集RefNLI,打破传统假设。
  • 模型在跨上下文场景下错误判断矛盾率达80%以上,误判蕴含超50%。
  • 揭示人类标注分歧部分源于参考不明确,适合关注NLI鲁棒性的研究者。

我们重新审视自然语言推理(NLI)任务中的参考确定性(RD)假设,即前提与假设被认为指向同一语境。尽管该假设便于构建新数据集,但当前仅基于此假设训练的NLI模型在事实验证等下游任务中表现不佳,因输入的前提与假设可能指向不同上下文。为揭示此现象在真实场景的影响,我们引入RefNLI——一个用于检测NLI例句中参考歧义的诊断基准。在RefNLI中,前提从知识源(如Wikipedia)检索,不一定与假设共享同一上下文。实验表明,微调后的NLI模型与少样本提示的大模型均无法识别上下文错配,导致超过80%的错误矛盾判断和超过50%的错误蕴含预测。我们发现,参考歧义的存在部分解释了人类标注中的内在分歧,并为RD假设对数据构建过程的影响提供了洞见。

原文摘要 · Abstract (English)

We revisit the reference determinacy (RD) assumption in the task of natural language inference (NLI), i.e., the premise and hypothesis are assumed to refer to the same context when human raters annotate a label. While RD is a practical assumption for constructing a new NLI dataset, we observe that current NLI models, which are typically trained solely on hypothesis-premise pairs created with the RD assumption, fail in downstream applications such as fact verification, where the input premise and hypothesis may refer to different contexts. To highlight the impact of this phenomenon in real-world use cases, we introduce RefNLI, a diagnostic benchmark for identifying reference ambiguity in NLI examples. In RefNLI, the premise is retrieved from a knowledge source (i.e., Wikipedia) and does not necessarily refer to the same context as the hypothesis. With RefNLI, we demonstrate that finetuned NLI models and few-shot prompted LLMs both fail to recognize context mismatch, leading to over 80% false contradiction and over 50% entailment predictions. We discover that the existence of reference ambiguity in NLI examples can in part explain the inherent human disagreements in NLI and provide insight into how the RD assumption impacts the NLI dataset creation process.

自然语言推理参考歧义模型评估事实验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。