arXiv:2504.04110cs.AIcs.CL2025-04ACL被引 17

用大模型联合推理材料与形式逻辑,提升论证可信度。

PEIRCE: Unifying Material and Formal Reasoning via LLM-Driven Neuro-Symbolic Refinement

  • 通过大模型生成自然语言与形式化表达的候选解
  • 结合符号证明器和软评估器迭代优化答案质量
  • 适合需要严谨推理与自然表达的场景

人工智能中一个持续的挑战是如何有效整合材料推理(关注论证的合理性与上下文相关性)与形式推理(关注逻辑与结构有效性)。大型语言模型凭借在海量文本上的预训练,具备强大的材料推理能力,但其推理常缺乏形式严谨性与可验证性。同时,语言模型的语言能力使其成为自然语言与形式语言之间的潜在桥梁,为融合两种推理模式提供新机遇。本文提出PEIRCE,一种基于大模型驱动的神经符号框架,通过迭代的猜想-批评过程统一材料与形式推理。在此框架中,大模型生成自然语言与形式语言下的候选解,再通过外部批判模型进行评估与精炼。这些批判模型包括符号证明器(用于检验形式有效性)以及软评估器(衡量生成论证在语言学与认识论维度的质量,如合理性、连贯性、简洁性)。尽管PEIRCE是通用框架,我们将其应用于自然语言解释生成任务,该任务天然要求材料充分性与形式正确性。

原文摘要 · Abstract (English)

A persistent challenge in AI is the effective integration of material and formal inference - the former concerning the plausibility and contextual relevance of arguments, while the latter focusing on their logical and structural validity. Large Language Models (LLMs), by virtue of their extensive pre-training on large textual corpora, exhibit strong capabilities in material inference. However, their reasoning often lacks formal rigour and verifiability. At the same time, LLMs' linguistic competence positions them as a promising bridge between natural and formal languages, opening up new opportunities for combining these two modes of reasoning. In this paper, we introduce PEIRCE, a neuro-symbolic framework designed to unify material and formal inference through an iterative conjecture-criticism process. Within this framework, LLMs play the central role of generating candidate solutions in natural and formal languages, which are then evaluated and refined via interaction with external critique models. These critiques include symbolic provers, which assess formal validity, as well as soft evaluators that measure the quality of the generated arguments along linguistic and epistemic dimensions such as plausibility, coherence, and parsimony. While PEIRCE is a general-purpose framework, we demonstrate its capabilities in the domain of natural language explanation generation - a setting that inherently demands both material adequacy and formal correctness.

神经符号大模型推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。