用生成流网络微调大模型,提升形式化推理能力。
Proof Flow: Preliminary Study on Generative Flow Network Language Model Tuning for Formal Reasoning
- 用生成流网络替代传统强化学习,更好探索推理路径。
- 在Lean形式证明任务中,模型搜索能力显著提升。
- 适合关注推理效率与模型泛化的研究者。
推理是解决新颖复杂问题的基础。尽管系统2式推理框架已取得进展,但开放模型仍难以应对高复杂度问题。本文探索生成流网络(GFlowNets)作为大语言模型微调方法的潜力,以解锁高级推理能力。在神经定理证明(NTP)场景下,使用形式语言如Lean编写可确定性验证的证明。与常规模型强化学习过度依赖高奖励动作、探索不足不同,GFlowNets能有效采样组合对象,增强泛化并保持多样假设。初步结果表明,该方法在搜索场景中显著提升模型表现,契合当前向推理时计算扩展和“慢思考”范式的转变。
原文摘要 · Abstract (English)
Reasoning is a fundamental substrate for solving novel and complex problems. Deliberate efforts in learning and developing frameworks around System 2 reasoning have made great strides, yet problems of sufficient complexity remain largely out of reach for open models. To address this gap, we examine the potential of Generative Flow Networks as a fine-tuning method for LLMs to unlock advanced reasoning capabilities. In this paper, we present a proof of concept in the domain of formal reasoning, specifically in the Neural Theorem Proving (NTP) setting, where proofs specified in a formal language such as Lean can be deterministically and objectively verified. Unlike classical reward-maximization reinforcement learning, which frequently over-exploits high-reward actions and fails to effectively explore the state space, GFlowNets have emerged as a promising approach for sampling compositional objects, improving generalization, and enabling models to maintain diverse hypotheses. Our early results demonstrate GFlowNet fine-tuning's potential for enhancing model performance in a search setting, which is especially relevant given the paradigm shift towards inference time compute scaling and "thinking slowly."
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。