用细粒度评估提升搜索推理模型的中间步骤质量
SRR-Judge: Step-Level Rating and Refinement for Enhancing Search-Integrated Reasoning in Search Agents
- 引入SRR-Judge框架,对推理与搜索每一步进行精准评分
- 在多个深度搜索基准上实现超过10%的准确率提升
- 适合需要高可靠性中间决策评估的研究者与开发者
近期基于大推理模型(LRMs)的深度搜索代理通过迭代规划、执行与证据收集,在复杂问题问答中表现优异,这一能力称为搜索集成推理。然而,主流方法仅依赖结果层面的监督,忽视了中间思考与动作的质量。本文提出SRR-Judge框架,实现对推理与搜索步骤的可靠细粒度评估。将其融入改进的ReAct风格“评分-修正”流程,提供精细化指导并支持高效后训练标注。利用SRR标注数据,采用迭代拒绝采样微调方法增强基础代理的深度搜索能力。实证表明,SRR-Judge的评估结果比更大模型如DeepSeek-V3.1更可靠,其评分与最终答案正确性高度相关。将策略对齐于SRR-Judge标注轨迹后,在多个挑战性深度搜索基准上实现了平均绝对pass@1提升超过10%。
原文摘要 · Abstract (English)
Recent deep search agents built on large reasoning models (LRMs) excel at complex question answering by iteratively planning, acting, and gathering evidence, a capability known as search-integrated reasoning. However, mainstream approaches often train this ability using only outcome-based supervision, neglecting the quality of intermediate thoughts and actions. We introduce SRR-Judge, a framework for reliable step-level assessment of reasoning and search actions. Integrated into a modified ReAct-style rate-and-refine workflow, SRR-Judge provides fine-grained guidance for search-integrated reasoning and enables efficient post-training annotation. Using SRR-annotated data, we apply an iterative rejection sampling fine-tuning procedure to enhance the deep search capability of the base agent. Empirically, SRR-Judge delivers more reliable step-level evaluations than much larger models such as DeepSeek-V3.1, with its ratings showing strong correlation with final answer correctness. Moreover, aligning the policy with SRR-Judge annotated trajectories leads to substantial performance gains, yielding over a 10 percent average absolute pass@1 improvement across challenging deep search benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。