arXiv:2501.16961cs.AI2025-01IJCAI被引 12

用实例验证提升大模型逻辑推理的准确性与可靠性。

Instantiation-based Formalization of Logical Reasoning Tasks using Language Models and Logical Solvers

  • 通过生成具体实例并让逻辑求解器验证,自动构建精确的逻辑形式化。
  • 在开放基准上实现比现有方法更高的推理准确率,验证精度接近完美。
  • 适合追求高可信推理的AI系统开发者,减少人工校验负担。

大型语言模型的推理鲁棒性仍是重大挑战,制约其实际应用。我们提出语义自验证(SSV),解决将自然语言推理问题转化为逻辑求解器可处理形式语言的核心难题。SSV采用基于一致性的方法,利用模型生成的具体实例经求解器验证,获得强抽象形式化。该方法显著提升整体推理准确率;更关键的是,在开放推理基准上实现了近乎完美的验证精度,覆盖广泛场景。我们提出‘近确定性推理’新范式,大幅降低多数情况下的手动验证需求,推动更可靠、自主的AI推理系统发展。

原文摘要 · Abstract (English)

Robustness of reasoning remains a significant challenge for large language models, and addressing it is essential for the practical applicability of AI-driven reasoning systems. We introduce Semantic Self-Verification (SSV), a novel approach that addresses the key challenge in combining language models with the rigor of logical solvers: to accurately formulate the reasoning problem from natural language to the formal language of the solver. SSV uses a consistency-based approach to produce strong abstract formalizations of problems using concrete instantiations that are generated by the model and verified by the solver. In addition to significantly advancing the overall reasoning accuracy over the state-of-the-art, a key novelty that this approach presents is a feature of verification that has near-perfect precision over a significant coverage of cases, as we demonstrate on open reasoning benchmarks. We propose such *near-certain reasoning* as a new approach to reduce the need for manual verification in many cases, taking us closer to more dependable and autonomous AI reasoning systems.

逻辑推理形式化自验证大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。