arXiv:2604.16004cs.CLcs.AI2026-04ACL被引 8

用双向智能体增强大模型推理的可信度,解决错误传播问题。

AgentV-RL: Scaling Reward Modeling with Agentic Verifier

论文配图:AgentV-RL: Scaling Reward Modeling with Agentic Verifier
图 1 · 摘自论文原文
  • 设计正反向双智能体,双向验证推理链可靠性。
  • 40亿参数模型比现有最优方法高25.2%准确率。
  • 适合需要高可信推理的复杂任务场景。

验证器通过测试时扩展(TTS)提升大模型推理能力,但在复杂领域面临挑战:错误推理会引发虚假通过,且缺乏外部依据导致在计算或知识密集型任务中不可靠。为此,我们提出Agentic Verifier,将奖励建模转化为多轮、工具增强的协商式过程。引入正向与反向智能体:前者从前提推导结论,后者从结论回溯前提,实现全面、可靠、可解释的评估。为促进实际部署,我们提出AgentV-RL,通过主动探索和强化学习,让验证器自主穿插使用工具与内部推理。大量实验表明,该框架在并行与串行TTS下均表现稳定提升,其40亿参数版本较当前最优ORM高出25.2%,展现出面向智能体奖励建模的潜力。

原文摘要 · Abstract (English)

Verifiers have been demonstrated to enhance LLM reasoning via test-time scaling (TTS). Yet, they face significant challenges in complex domains. Error propagation from incorrect intermediate reasoning can lead to false positives for seemingly plausible solutions, while lacking external grounding makes verifiers unreliable on computation or knowledge-intensive tasks. To address these challenges, we propose Agentic Verifier, a framework that transforms reward modeling into a multi-turn, tool-augmented deliberative process. We introduce complementary forward and backward agents: one traces solutions from premises to conclusions, while the other re-checks conclusions against their underlying premises. This bidirectional process enables a comprehensive, reliable, and interpretable assessment of solutions. To facilitate practical deployment, we propose AgentV-RL. Through proactive exploration and reinforcement learning, the verifier autonomously interleaves tool-use with internal reasoning. Extensive experiments show that Agentic Verifier yields consistent performance gains under both parallel and sequential TTS. Notably, our 4B variant surpasses state-of-the-art ORMs by 25.2%, positioning it as a promising paradigm for agentic reward modeling.

智能体推理验证奖励建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。