arXiv:2505.22650cs.LG2025-05NeurIPS被引 5

训练可验证自然语言推理步骤的模型,提升复杂问题求解可靠性

On Learning Verifiers and Implications to Chain-of-Thought Reasoning

  • 构建形式化学习框架,评估自然语言推理的正确性
  • 给出不同验证强度下的样本复杂度上界与不可行性结论
  • 为链式思维推理提供可信验证机制,适合逻辑与数学推理研究者

链式思维推理已成为解决复杂数学与逻辑问题的有效方法,但常因错误或无根据的推断偏离正轨。形式化数学推理可通过形式验证器检查,但当前大模型尚无法以形式方式解决复杂问题,甚至将非形式问题陈述形式化也极具挑战。为此,本文研究如何学习可靠的自然语言链式思维推理验证器:给定问题陈述与逐步解答,验证器输出[是]若所有推理步骤均有效,否则输出[否]。本文提出一个形式化的PAC学习框架,分析多个不同强度的验证目标,给出满足这些目标的验证器的学习样本复杂度上界,并证明在无额外假设下,其他自然验证目标不可学习。

原文摘要 · Abstract (English)

Chain-of-Thought reasoning has emerged as a powerful approach for solving complex mathematical and logical problems. However, it can often veer off track through incorrect or unsubstantiated inferences. Formal mathematical reasoning, which can be checked with a formal verifier, is one approach to addressing this issue. However, currently LLMs are simply not good enough to solve complex problems in a formal way, and even just formalizing an informal problem statement can be challenging. Motivated by this fact, in this work we consider the problem of learning reliable verifiers for natural language Chain-of-Thought reasoning. That is, given a problem statement and step-by-step solution in natural language, the aim of the verifier is to output [Yes] if the reasoning steps in the solution are all valid, and [No] otherwise. In this work we give a formal PAC-learning framework for studying this problem. We propose and analyze several natural verification goals, at different levels of strength, in this framework. We provide sample complexity upper-bounds for learning verifiers satisfying these goals, as well as lower-bound and impossibility results for learning other natural verification objectives without additional assumptions.

链式思维推理验证形式化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。