arXiv:2502.16757cs.CL2025-02ACL被引 5

让自然语言推理的逻辑表达更准确,提升自动证明的保真率。

Entailment-Preserving First-order Logic Representations in Natural Language Entailment

  • 设计新任务与评估指标,直接优化逻辑表达对蕴含关系的保持能力。
  • 提出迭代学习排序方法,在三个数据集上使保真率提升1.8%-2.7%。
  • 有效减少逻辑表达的随意性,适用于多种推理类型和跨领域数据。

一阶逻辑(FOL)可表征自然语言句子间的逻辑蕴含语义,但利用FOL判断自然语言蕴含仍具挑战。为此,我们提出蕴含保持型一阶逻辑表示任务(EPF),并引入无参考评估指标——蕴含保持率(EPR)家族。在EPF中,需从多前提自然语言蕴含数据(如EntailmentBank)生成FOL表示,使自动证明器结果保持原始蕴含标签。实验表明,现有NL到FOL转换方法在该任务上表现不佳。为此,我们提出一种专用于EPF的训练方法:迭代学习排序,通过新型评分函数与学习排序目标直接优化模型的EPR得分。该方法在三个数据集上实现1.8%-2.7%的EPR提升,以及17.4%-20.6%的EPR@16增长。进一步分析显示,该方法有效抑制了FOL表示的任意性,降低谓词签名多样性,并在多种推理类型及域外数据上保持强性能。

原文摘要 · Abstract (English)

First-order logic (FOL) can represent the logical entailment semantics of natural language (NL) sentences, but determining natural language entailment using FOL remains a challenge. To address this, we propose the Entailment-Preserving FOL representations (EPF) task and introduce reference-free evaluation metrics for EPF, the Entailment-Preserving Rate (EPR) family. In EPF, one should generate FOL representations from multi-premise natural language entailment data (e.g. EntailmentBank) so that the automatic prover's result preserves the entailment labels. Experiments show that existing methods for NL-to-FOL translation struggle in EPF. To this extent, we propose a training method specialized for the task, iterative learning-to-rank, which directly optimizes the model's EPR score through a novel scoring function and a learning-to-rank objective. Our method achieves a 1.8-2.7% improvement in EPR and a 17.4-20.6% increase in EPR@16 compared to diverse baselines in three datasets. Further analyses reveal that iterative learning-to-rank effectively suppresses the arbitrariness of FOL representation by reducing the diversity of predicate signatures, and maintains strong performance across diverse inference types and out-of-domain data.

逻辑推理自然语言一阶逻辑评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。