arXiv:2509.22338cs.CLcs.AI2025-09中稿 · the International …被引 4

用微调大模型提升自然语言转一阶逻辑的准确率,效果优于GPT-4o。

Advancing Natural Language Formalization to First Order Logic with Fine-tuned LLMs

  • 微调T5类模型,结合谓词条件与词汇扩展提升翻译能力
  • 在MALLS数据集上达到70%精确匹配准确率,优于GPT-4o
  • 适用于需要高精度逻辑形式化的研究者或系统开发者

自动化将自然语言翻译为一阶逻辑(FOL)对知识表示与形式化方法至关重要,但依然具有挑战性。本文系统评估了微调大模型在该任务中的表现,对比了编码器-解码器与仅解码器架构及训练策略。基于MALLS和Willow数据集,探索了词汇扩展、谓词条件化与多语言训练等技术,并引入精确匹配、逻辑等价与谓词对齐等评估指标。微调后的Flan-T5-XXL在提供谓词列表的情况下达到70%准确率,超越GPT-4o以及具备思维链推理能力的DeepSeek-R1-0528模型,也优于符号系统ccg2lambda。关键发现包括:(1) 谓词可用性使性能提升15-20%;(2) T5模型优于更大规模的仅解码器模型;(3) 模型可在未见逻辑论据(如FOLIO数据集)上泛化,无需专门训练。尽管结构化逻辑翻译表现稳健,谓词提取仍是主要瓶颈。

原文摘要 · Abstract (English)

Automating the translation of natural language to first-order logic (FOL) is crucial for knowledge representation and formal methods, yet remains challenging. We present a systematic evaluation of fine-tuned LLMs for this task, comparing architectures (encoder-decoder vs. decoder-only) and training strategies. Using the MALLS and Willow datasets, we explore techniques like vocabulary extension, predicate conditioning, and multilingual training, introducing metrics for exact match, logical equivalence, and predicate alignment. Our fine-tuned Flan-T5-XXL achieves 70% accuracy with predicate lists, outperforming GPT-4o and even the DeepSeek-R1-0528 model with CoT reasoning ability as well as symbolic systems like ccg2lambda. Key findings show: (1) predicate availability boosts performance by 15-20%, (2) T5 models surpass larger decoder-only LLMs, and (3) models generalize to unseen logical arguments (FOLIO dataset) without specific training. While structural logic translation proves robust, predicate extraction emerges as the main bottleneck.

自然语言逻辑形式化大模型微调一阶逻辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。