arXiv:2609.02417cs.LGcs.AI2026-09

多轮智能体强化学习中,均匀分配奖励比精准定位关键步骤更有效。

Coverage, Not Targeting: A Structural Regime in Multi-Turn Agent Credit Assignment

  • 提出验证信息密度V_d衡量奖励分配是否应聚焦特定步骤
  • 实验显示均匀奖励在多数场景优于集中奖励,即使集中在正确步骤也无效
  • 适用于研究智能体信用分配机制的学者与开发多步任务系统者

多轮智能体强化学习常将信用分配视为定位问题:给定可验证的最终奖励,逐轮方法试图锁定关键回合。本文识别出决定该策略是否有效的结构性指标——验证信息密度 V_d = k/C(验证器暴露的因果链比例),发现终态验证器处于低V_d区域,此时聚焦定位是错误方向。在tau^2-bench上的受控共享回放实验表明,连续密集奖励普遍优于稀疏二值奖励(5个种子中有4个表现更差);而将相同优势集中在进展回合或随机回合均导致性能下降,说明聚焦是次级效应。其机制在于覆盖性:终态验证将可观测信号压缩至单一轮次(98%回放中k=1),但成功需5-8步前置工具调用。合成相变边界显示临界点V_d*≈0.8,而tau^2-bench实测为~0.15,BFCL V3为~0.4;在BFCL上均匀奖励仍胜出,匹配集中度的随机控制组在8/8种子中为负。该现象在ToolACE-2-8B模型族中重复出现(32个预注册种子中Delta=-0.048;独立20种子复制亦显著)。预注册的匹配预算广度扫描显示单调剂量响应,仅当全链覆盖时缺陷消失,且基于未来回报的臂能达到全覆盖平衡。均匀重分配是零信息覆盖基准,任何聚焦主张必须超越此基准。本文贡献了匹配集中度的随机控制方案,作为验证聚焦假说的必要门槛。

原文摘要 · Abstract (English)

Multi-turn agentic RL increasingly treats credit assignment as a targeting problem: given a terminal verifiable reward, per-turn methods localize credit onto the turns that mattered. We identify the structural quantity that predicts when this is the right move, the verifier information density V_d = k/C (the fraction of an agent's C-step causal chain whose per-turn correctness the verifier exposes), and show that terminal-state verifiers sit deep in a low-V_d regime where targeting is the wrong axis. In controlled shared-rollout comparisons on tau^2-bench that separate reward density from credit geometry, a continuous dense reward spread uniformly beats the sparse binary outcome reward (net-harmful on 4/5 seeds), while concentrating the same advantage on progress turns or on random turns is equally harmful: targeting is second-order. The mechanism is coverage: terminal-state verification collapses the observable signal to a single final-write turn (k=1 in 98% of rollouts) while success requires a 5-8 step chain of prerequisite tool calls. A synthetic phase boundary places the crossover at V_d* ~ 0.8, whereas measured V_d is ~0.15 on tau^2-bench and ~0.4 on BFCL V3; uniform also wins on BFCL, where a matched-concentration shuffled control is negative on 8/8 seeds. The effect reproduces across model families on ToolACE-2-8B (Delta = -0.048 over 32 pre-registered seeds; an independent 20-seed replication is itself significant), and a pre-registered matched-budget breadth sweep traces a monotone dose-response whose deficit vanishes only at full chain coverage, with a reward-to-go arm reaching full-coverage parity. Uniform redistribution is the zero-information coverage default that per-turn schemes must beat; we contribute the matched-concentration shuffled control that any targeting claim should clear.

强化学习信用分配多步推理智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。