arXiv:2604.14877cs.LG2026-04被引 3

强化学习让大模型代理真正突破能力边界,尤其在复杂任务中

Does RL Expand the Capability Boundary of LLM Agents? A PASS@(k,T) Analysis

论文配图:Does RL Expand the Capability Boundary of LLM Agents? A PASS@(k,T) Analysis
图 1 · 摘自论文原文
  • 提出二维指标PASS@(k,T),区分能力扩展与效率提升
  • 在组合式任务中,强化学习使通过率曲线持续高于基线,差距随预算增大
  • 能力提升源于自主探索带来的信息整合优化,非单纯可靠性增强

强化学习是否真正拓展了大模型代理的能力边界,还是仅提升其可靠性?针对静态推理,已有研究认为两者在大采样量k下趋于收敛。本文考察代理使用工具的动态交互场景,其中T轮交互允许生成复合策略,而重采样无法恢复此类策略。我们引入PASS@(k,T)——一个同时变化采样预算k与交互深度T的二维指标,以分离能力扩展与效率改进。主要发现:与静态推理相反,在工具使用任务中,强化学习确实显著扩展了能力边界——强化学习代理的通过率曲线始终高于基线,并在大k下差距扩大而非收敛。该扩展仅出现在组合性、序列化信息获取任务中;简单任务上表现如过往研究预测。在相同训练数据下,监督微调反而压缩了该边界,表明自导向探索是关键因果因素。机制分析显示,强化学习将基线策略分布重加权至下游推理更易出正确答案的子集,且改进集中于信息整合环节。这些结果调和了关于强化学习对大模型的乐观与悲观观点:二者均成立,但适用于不同任务类型。

原文摘要 · Abstract (English)

Does reinforcement learning genuinely expand what LLM agents can do, or merely make them more reliable? For static reasoning, recent work answers the second: base and RL pass@k curves converge at large k. We ask whether this holds for agentic tool use, where T rounds of interaction enable compositional strategies that re-sampling cannot recover. We introduce PASS@(k,T), a two-dimensional metric that jointly varies sampling budget k and interaction depth T, separating capability expansion from efficiency improvement. Our main finding is that, contrary to the static-reasoning result, tool-use RL genuinely enlarges the capability boundary: the RL agent's pass-curve pulls above the base model's and the gap widens at large k rather than converging. The expansion is specific to compositional, sequential information gathering; on simpler tasks RL behaves as prior work predicts. Under matched training data, supervised fine-tuning regresses the boundary on the same compositional tasks, isolating self-directed exploration as the causal factor. Mechanism analysis shows RL reweights the base strategy distribution toward the subset whose downstream reasoning more often yields a correct answer, with the improvement concentrated on how the agent integrates retrieved information. These results reconcile optimistic and pessimistic readings of RL for LLMs: both are correct, on different task types.

强化学习大模型代理能力边界工具使用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。