让AI研究代理在推理时自我纠错,逐步提升答案质量。
Inference-Time Scaling of Verification: Self-Evolving Deep Research Agents via Test-Time Rubric-Guided Verification
- 用自建评分标准自动检验生成结果,反馈优化
- 在多个数据集上比基线提升8%-11%准确率
- 适合想提升AI推理能力的研究者和开发者
深度研究代理(DRAs)正推动自动化知识发现与问题求解。现有方法多通过后训练增强策略能力,本文提出新范式:通过精心设计的评分标准,在推理阶段迭代验证策略输出,实现自我演化。基于自动构建的DRA失败分类体系,将失败分为五类十三子类,设计出可指导评估的评分标准。提出DeepVerifier,一种基于评分标准的结果奖励验证器,利用验证与生成的不对称性,在元评估F1分数上优于基线12%-48%。该模块可插拔式集成于测试阶段,生成细粒度反馈供代理迭代改进,无需额外训练。在具备强大闭源大模型支持下,使GAIA和XBench-DeepSearch挑战子集准确率提升8%-11%。为推动开源发展,发布DeepVerifier-4K,包含4,646条高质量代理步骤数据,强调反思与自省,助力开源模型建立稳健验证能力。
原文摘要 · Abstract (English)
Recent advances in Deep Research Agents (DRAs) are transforming automated knowledge discovery and problem-solving. While the majority of existing efforts focus on enhancing policy capabilities via post-training, we propose an alternative paradigm: self-evolving the agent's ability by iteratively verifying the policy model's outputs, guided by meticulously crafted rubrics. This approach gives rise to the inference-time scaling of verification, wherein an agent self-improves by evaluating its generated answers to produce iterative feedback and refinements. We derive the rubrics based on an automatically constructed DRA Failure Taxonomy, which systematically classifies agent failures into five major categories and thirteen sub-categories. We present DeepVerifier, a rubrics-based outcome reward verifier that leverages the asymmetry of verification and outperforms vanilla agent-as-judge and LLM judge baselines by 12%-48% in meta-evaluation F1 score. To enable practical self-evolution, DeepVerifier integrates as a plug-and-play module during test-time inference. The verifier produces detailed rubric-based feedback, which is fed back to the agent for iterative bootstrapping, refining responses without additional training. This test-time scaling delivers 8%-11% accuracy gains on challenging subsets of GAIA and XBench-DeepSearch when powered by capable closed-source LLMs. Finally, to support open-source advancement, we release DeepVerifier-4K, a curated supervised fine-tuning dataset of 4,646 high-quality agent steps focused on DRA verification. These examples emphasize reflection and self-critique, enabling open models to develop robust verification capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。