arXiv:2502.08080cs.CL2025-02NAACL被引 10

拆解假设为基本命题,揭示大模型在自然语言推理中的逻辑漏洞。

NLI under the Microscope: What Atomic Hypothesis Decomposition Reveals

  • 将假设分解为原子命题,细粒度分析模型推理过程。
  • 发现大模型在原子层面仍存在逻辑不一致问题。
  • 提出新指标衡量模型对同一事实的推理一致性,适合评估模型可靠性。

将文本分解为原子命题是一种灵活的框架,可用于更深入地分析输入与输出文本。我们将在传统自然语言推理(NLI)和可反驳性自然语言推理(defeasible NLI)任务中,对假设进行原子分解,形成细粒度的子问题,即模型在解决整体问题时必须权衡的原子推断。这些子问题有助于进一步理解NLI与可反驳推理的结构,探测模型在不同推断上的一致性与理解程度,并衡量基准数据集示例的多样性。结果表明,大型语言模型在原子层面的NLI与可反驳性NLI子问题上仍存在逻辑不一致。最后,我们识别出可反驳性NLI示例中起关键作用的原子子问题,并提出一种测量模型推理性一致性的方法,该指标旨在捕捉模型在不同上下文中对同一事实做出一致正确或错误预测的程度。

原文摘要 · Abstract (English)

Decomposition of text into atomic propositions is a flexible framework allowing for the closer inspection of input and output text. We use atomic decomposition of hypotheses in two natural language reasoning tasks, traditional NLI and defeasible NLI, to form atomic sub-problems, or granular inferences that models must weigh when solving the overall problem. These atomic sub-problems serve as a tool to further understand the structure of both NLI and defeasible reasoning, probe a model's consistency and understanding of different inferences, and measure the diversity of examples in benchmark datasets. Our results indicate that LLMs still struggle with logical consistency on atomic NLI and defeasible NLI sub-problems. Lastly, we identify critical atomic sub-problems of defeasible NLI examples, or those that most contribute to the overall label, and propose a method to measure the inferential consistency of a model, a metric designed to capture the degree to which a model makes consistently correct or incorrect predictions about the same fact under different contexts.

自然语言推理逻辑一致性大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。