arXiv:2605.26352cs.CL2026-05

让检索互动自动生成可信的训练信号,提升推理型智能体表现。

RICE-PO: Turning Retrieval Interactions into Credit Signals for Reasoning Agents

论文配图:RICE-PO: Turning Retrieval Interactions into Credit Signals for Reasoning Agents
图 1 · 摘自论文原文
  • 通过高不确定性的操作作为锚点,构建局部反事实分支评估
  • 在BRIGHT和BEIR数据集上超越提示型与群体强化学习基线
  • 无需评判器,利用交互结构自身提供可信赖的监督信号

检索正从单次匹配转向交互式推理,语言智能体迭代检查证据、重写查询并再次搜索。训练这类智能体面临信用分配难题:可执行动作(如查询或摘要)可被检索器直接评估,而隐含的推理步骤不可观测,仅通过影响后续可执行动作间接体现。这种不对称性导致结果级奖励不可靠——相同最终奖励可能错误归因于未实际影响检索成功的推理步骤。我们提出RICE-PO,一种无评判器的策略优化框架,将检索互动转化为局部学习信号。RICE-PO选取高不确定性可执行动作作为锚点,利用检索指标评估局部反事实分支,并仅当推理到动作的影响强且未来残差效应稳定时,才向隐含推理步骤传播信用。在BRIGHT和BEIR数据集上,RICE-PO在相同检索器设置下持续优于提示型智能体和群体强化学习基线。结果表明,智能体-环境交互结构本身即可为基于推理的检索智能体提供有效监督。

原文摘要 · Abstract (English)

Retrieval is increasingly moving from one-shot matching toward interactive reasoning, where language agents iteratively inspect evidence, reformulate queries, and search again. Training such agents raises a credit-assignment challenge: executable actions such as queries or summaries can be directly evaluated by the retriever, while latent reasoning steps are not directly observable and only affect future executable actions. This asymmetry makes outcome-level reward assignment unreliable, as the same final reward may credit reasoning steps that did not actually shape retrieval success. We propose RICE-PO, a critic-free policy optimization framework that converts retrieval interactions into localized learning signals. RICE-PO selects high-uncertainty executable actions as anchors, evaluates local counterfactual branches using retrieval metrics, and propagates credit to latent reasoning steps only when reasoning-to-action influence is strong and future residual effects are stable. On BRIGHT and BEIR, RICE-PO consistently outperforms prompt-based agents and group-based RL baselines under the same retriever setting. These results show that the structure of agent-environment interaction itself can provide useful supervision for training reasoning-based retrieval agents.

推理增强检索学习强化学习智能体训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。