低决策密度导致多轮智能体训练信号稀释,影响学习效率。
Drowning in Routine: Signal Dilution in Multi-Turn Agent Training

- 用决策密度ρ衡量关键动作占比,揭示信号稀释根源
- 当ρ趋近0时,训练步数差距显著扩大,验证理论预测
- 高决策密度下可省去价值网络,降低训练成本
多轮智能体在执行中混合关键决策与常规操作:部分动作影响最终回报分布,而其他动作虽必要但奖励等价。轨迹级信用分配的代价并非仅由长时序引起,而是由决策密度ρ(影响回报的动作比例)决定。当ρ较低时,常规轮次造成信号稀释:增加梯度方差但不提升期望信号,使转换单位级到轨迹级的信噪比下降至ρ^{-1/2},前提是价值函数误差可控。分析还识别出互补情形:高决策密度下,无需价值网络的轨迹级方法仍具竞争力。在可精确调节ρ的受控环境中,实验验证了该标度关系,$R^2 = 0.999$,且随着ρ→0,训练步数差距显著扩大。
原文摘要 · Abstract (English)
Multi-turn agents interleave consequential decisions with routine execution: some actions change the downstream return distribution, while others are necessary but reward-equivalent. The cost of trajectory-level credit assignment, often attributed to long horizons, is in fact governed by decision density $ρ$: the fraction of turns whose actions affect the return. When decision density is low, routine turns create signal dilution: they add gradient variance to trajectory-level estimators such as GRPO without adding expected signal. Under explicit assumptions, the resulting turn-level to trajectory-level signal-to-noise ratio scales as $ρ^{-1/2}$, provided critic error remains controlled. The same analysis identifies the complementary regime: at high decision density, trajectory-level methods can remain competitive while avoiding the cost of a critic. In a controlled environment where $ρ$ is exactly tunable, the predicted scaling is recovered with $R^2 = 0.999$, and the training-step gap widens significantly as $ρ\to 0$.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。