构建新数据集评估模型对条件句预设的推理能力
Let's CONFER: A Dataset for Evaluating Natural Language Inference Models on CONditional InFERence and Presupposition
- 提出CONFER数据集,专测条件句中的预设推理
- 现有NLI模型在条件句预设任务上表现不佳
- 大模型如GPT-4o在零样本下仍具较好推理能力
自然语言推理(NLI)旨在判断句子对之间是否存在蕴含、矛盾或中立关系。尽管当前NLI模型在多数任务上表现良好,但其对细粒度语用推理,尤其是条件句中的预设推理能力仍缺乏充分研究。本文提出CONFER数据集,用于评估NLI模型在条件句推理中的表现。我们测试了四种NLI模型(含两个预训练模型)在条件推理上的泛化能力,并评估了GPT-4o、LLaMA、Gemma和DeepSeek-R1等大语言模型在零样本与少样本提示下的表现,分析其在有无上下文条件下推断预设的能力。结果表明,现有NLI模型在条件句预设推理上存在显著不足,且在现有NLI数据集上微调并不能提升其性能。
原文摘要 · Abstract (English)
Natural Language Inference (NLI) is the task of determining whether a sentence pair represents entailment, contradiction, or a neutral relationship. While NLI models perform well on many inference tasks, their ability to handle fine-grained pragmatic inferences, particularly presupposition in conditionals, remains underexplored. In this study, we introduce CONFER, a novel dataset designed to evaluate how NLI models process inference in conditional sentences. We assess the performance of four NLI models, including two pre-trained models, to examine their generalization to conditional reasoning. Additionally, we evaluate Large Language Models (LLMs), including GPT-4o, LLaMA, Gemma, and DeepSeek-R1, in zero-shot and few-shot prompting settings to analyze their ability to infer presuppositions with and without prior context. Our findings indicate that NLI models struggle with presuppositional reasoning in conditionals, and fine-tuning on existing NLI datasets does not necessarily improve their performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。