arXiv:2509.19771cs.LGcs.AI2025-09

提出摩擦式Q学习,缓解离策略强化学习的外推误差问题。

Frictional Q-Learning

  • 用对比变分自编码器将支持动作建模为切向方向。
  • 发现价值敏感度存在各向异性,自然形成类似摩擦阈值的稳定条件。
  • 在连续控制基准测试中表现更稳定,适合高可靠性强化学习场景。

离策略强化学习在策略选择未充分支持的动作时会引发外推误差。本文借鉴静摩擦概念,将经验回放缓冲区视为低维动作流形,支持方向对应切向分量,法向分量则捕获主导的一阶外推误差。该分解揭示了价值敏感度的内在各向异性,自然引出类似摩擦阈值的稳定性条件。为此,我们提出摩擦式Q学习(Frictional Q-Learning),利用对比变分自编码器将支持动作编码为切向方向。在温和局部等距假设下,正交补空间的正交基对应于法向分量。在标准连续控制基准上的大量实验表明,该方法相比竞争基线展现出更稳健和稳定的性能。

原文摘要 · Abstract (English)

Off-policy reinforcement learning suffers from extrapolation errors when a learned policy selects actions that are weakly supported in the replay buffer. In this study, we address this issue by drawing an analogy to static friction. From this perspective, the replay buffer is represented as a smooth, low-dimensional action manifold, where the support directions correspond to the tangential component, while the normal component captures the dominant first-order extrapolation error. This decomposition reveals an intrinsic anisotropy in value sensitivity that naturally induces a stability condition analogous to a friction threshold. To mitigate deviations toward unsupported actions, we propose Frictional Q-Learning, an off-policy algorithm that encodes supported actions as tangent directions using a contrastive variational autoencoder. We further show that an orthonormal basis of the orthogonal complement corresponds to normal components under mild local isometry assumptions. Extensive empirical results on standard continuous-control benchmarks consistently demonstrate robust and stable performance compared with competitive baselines.

强化学习离策略动作支持稳定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。