arXiv:2602.10430cs.LGcs.AI2026-02

解决负样本更新过度排斥问题,让强化学习更稳定高效

Breaking the Curse of Repulsion: Remoteness-Aware Control of Negative Off-Policy Updates

  • 提出动态远近感知策略优化,按距离衰减远端负样本影响
  • 实验证明中间偏移能提升保留奖励,恢复收敛性
  • 适合做离线强化学习与策略优化的科研人员

离线策略优化利用历史行为数据,包括负优势样本以抑制已知失败动作。我们发现重复使用会将有益信号转化为过度排斥:当学习器远离历史负动作后,后续更新使该动作愈发遥远,但更新强度未必减弱。理论分析揭示了从稳定位移过渡到持续漂移、失去有限稳定平衡点的过程;可控强度扫描显示,适度偏移可提升保留奖励。相关相对坐标为高斯策略的平方标准化距离或分类策略的意外度。我们提出动态远近感知策略优化(DRPO),在近场保持负更新不变,在远场指数衰减其尾部。DRPO 对任意固定负正样本比均实现最终向内漂移,给出明确极限半径,并将分类策略的支持抑制由指数衰减变为多项式衰减。外部诊断与干预表明,远近依赖的策略几何是负更新放大的根源,选择性截断可消除远场效应而不丢弃局部有效反馈。

原文摘要 · Abstract (English)

Off-policy policy optimization reuses historical behavior, including negative-advantage samples that suppress known failures. We show that repeated reuse can turn this useful signal into excessive repulsion: as the learner moves away from a historical negative action, subsequent updates make that action increasingly remote without necessarily reducing its update strength. Our aggregate theory characterizes the resulting transition from a stable displacement beyond the positive-only target to persistent drift and the loss of finite stable equilibria; controlled strength sweeps show that an intermediate displacement can improve held-out reward. The relevant learner-relative coordinate is squared standardized distance for Gaussian policies and surprisal for categorical policies. We introduce Dynamic Remoteness-Aware Policy Optimization (DRPO), which leaves the negative update unchanged in the near field and exponentially attenuates its remote tail. DRPO restores eventual inward Gaussian drift for every fixed finite negative-to-positive mass ratio, yields an explicit ultimate-bound radius, and changes categorical support suppression from exponential to polynomial probability decay. External diagnostics and controlled interventions isolate remoteness-dependent policy geometry as a source of negative-update amplification and show that selective tapering can remove its far-field effect without discarding useful local feedback.

强化学习离线策略策略优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。