用生成式方法解决连续控制中的信号延迟问题,兼具模型感知与决策能力。
State-Action Inpainting Diffuser for Continuous Control with Delay
- 将延迟控制建模为状态-动作补全任务,融合动力学先验与策略优化
- 在多个延迟控制基准上达到顶尖性能,且适用于在线与离线强化学习
- 提出新范式:生成式补全兼顾动态建模与直接决策,适合延迟敏感场景
信号延迟在连续控制与强化学习中造成交互与感知间的时序鸿沟。现有方法主要分为两类:无模型方法通过状态增强维持马尔可夫性,有模型方法则聚焦于通过动态建模推断潜在信念。本文提出状态-动作补全扩散器(SAID),融合动力学学习的归纳偏置与策略优化的直接决策能力。将问题建模为联合序列补全任务,隐式捕捉环境动态的同时直接生成一致的行动规划,在模型基与无模型范式交界处有效运作。该生成式框架可无缝应用于在线与离线强化学习。在多个带延迟的连续控制基准上的实验表明,SAID实现最先进且稳健的性能。本研究为延迟强化学习提供了新方法论。
原文摘要 · Abstract (English)
Signal delay poses a fundamental challenge in continuous control and reinforcement learning (RL) by introducing a temporal gap between interaction and perception. Current solutions have largely evolved along two distinct paradigms: model-free approaches which utilize state augmentation to preserve Markovian properties, and model-based methods which focus on inferring latent beliefs via dynamics modeling. In this paper, we bridge these perspectives by introducing State-Action Inpainting Diffuser (SAID), a framework that integrates the inductive bias of dynamics learning with the direct decision-making capability of policy optimization. By formulating the problem as a joint sequence inpainting task, SAID implicitly captures environmental dynamics while directly generating consistent plans, effectively operating at the intersection of model-based and model-free paradigms. Crucially, this generative formulation allows SAID to be seamlessly applied to both online and offline RL. Extensive experiments on delayed continuous control benchmarks demonstrate that SAID achieves state-of-the-art and robust performance. Our study suggests a new methodology to advance the field of RL with delay.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。