用有向无环图实现大模型推理的灵活验证
Graph of Verification: Structured Verification of LLM Reasoning with Directed Acyclic Graphs
- 通过可变粒度节点块,动态匹配不同结构的推理过程
- 在结构化与非结构化任务上均显著优于现有方法
- 适合需要高精度推理验证的研究者和开发者
大语言模型的多步推理验证面临关键挑战:整体性验证常忽略局部错误。逐步验证虽有潜力,但现有方法缺乏灵活性,难以适应从正式证明到自然语言叙述的多样化推理结构。为此,我们提出图式验证(Graph of Verification, GoV),一种可适配、多粒度的验证框架。其核心是灵活的“节点块”架构,能根据推理内容自动调整验证粒度——从形式化任务中的原子步骤,到自然语言叙述中的整段内容。该机制有效缓解了验证精度与鲁棒性之间的权衡。在结构化与松散结构的多个基准测试中,GoV均展现出卓越适应性,其自适应方法显著超越整体基线及当前最优分解式验证方法,为无需训练的推理验证树立了新标准。
原文摘要 · Abstract (English)
Verifying the complex and multi-step reasoning of Large Language Models (LLMs) is a critical challenge, as holistic methods often overlook localized flaws. Step-by-step validation is a promising alternative, yet existing methods are often rigid. They struggle to adapt to diverse reasoning structures, from formal proofs to informal natural language narratives. To address this adaptability gap, we propose the Graph of Verification (GoV), a novel framework for adaptable and multi-granular verification. GoV's core innovation is its flexible "node block" architecture. This mechanism allows GoV to adaptively adjust its verification granularity--from atomic steps for formal tasks to entire paragraphs for natural language--to match the native structure of the reasoning process. This flexibility allows GoV to resolve the fundamental trade-off between verification precision and robustness. Experiments on both well-structured and loosely-structured benchmarks demonstrate GoV's versatility. The results show that GoV's adaptive approach significantly outperforms both holistic baselines and other state-of-the-art decomposition-based methods, establishing a new standard for training-free reasoning verification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。