arXiv:2602.03647cs.AIcs.CL2026-02被引 10

通过协作式纠错提升语言模型搜索推理能力

Search-R2: Enhancing Search-Integrated Reasoning via Actor-Refiner Collaboration

  • 分角色设计生成与修正模块,实现精准干预
  • 在多跳问答任务中准确率显著超越现有方法
  • 适合需要高可靠推理的智能问答系统

搜索集成推理使语言代理能够超越静态参数化知识,主动查询外部信息。然而,通过强化学习训练这些代理受到多尺度信用分配问题的阻碍:现有方法通常依赖稀疏的轨迹级奖励,无法区分高质量推理与偶然猜测,导致冗余或误导性搜索行为。为此,我们提出Search-R2,一种新型的演员-修正者协作框架,通过针对性干预增强推理能力,且两个组件在训练中联合优化。我们的方法将生成过程分解为演员(产生初始推理轨迹)和元修正者(通过‘剪切并重生成’机制选择性诊断和修复错误步骤)。为提供细粒度监督,我们引入混合奖励设计,将结果正确性与量化检索证据信息密度的密集过程奖励结合。理论上,我们将演员-修正者交互形式化为平滑混合策略,证明选择性修正相比强基线有严格性能提升。在多种通用及多跳问答数据集上的大量实验表明,Search-R2始终优于强基线,在不同模型规模下均实现更优推理准确率,且开销极小。

原文摘要 · Abstract (English)

Search-integrated reasoning enables language agents to transcend static parametric knowledge by actively querying external sources. However, training these agents via reinforcement learning is hindered by the multi-scale credit assignment problem: existing methods typically rely on sparse, trajectory-level rewards that fail to distinguish between high-quality reasoning and fortuitous guesses, leading to redundant or misleading search behaviors. To address this, we propose Search-R2, a novel Actor-Refiner collaboration framework that enhances reasoning through targeted intervention, with both components jointly optimized during training. Our approach decomposes the generation process into an Actor, which produces initial reasoning trajectories, and a Meta-Refiner, which selectively diagnoses and repairs flawed steps via a 'cut-and-regenerate' mechanism. To provide fine-grained supervision, we introduce a hybrid reward design that couples outcome correctness with a dense process reward quantifying the information density of retrieved evidence. Theoretically, we formalize the Actor-Refiner interaction as a smoothed mixture policy, proving that selective correction yields strict performance gains over strong baselines. Extensive experiments across various general and multi-hop QA datasets demonstrate that Search-R2 consistently outperforms strong RAG and RL-based baselines across model scales, achieving superior reasoning accuracy with minimal overhead.

搜索推理强化学习协同优化问答系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。