用LSTM+强化学习缓解远程操控中的随机延迟问题
Residual Reinforcement Learning for Robot Teleoperation under Stochastic Delays

- 用LSTM重构延迟观测下的连续状态,提升控制稳定性
- 在Franka Panda上实现比现有方法更稳定、平滑的远程操控
- 适合高延迟波动场景下的机器人遥操作应用
遥操作中的随机通信延迟导致信号不连续,破坏控制稳定性并降低性能。传统强化学习因延迟观测难以应对,易产生高频抖动。为此,我们提出一种混合控制框架——延迟鲁棒强化学习,结合基于长短期记忆(LSTM)的状态估计算法与残差强化学习策略,使代理能够从延迟观测中重建平滑连续的状态,并学习残差力矩补偿策略,在跟踪精度与速度平滑性之间取得平衡。在Franka Panda机器人上的实验表明,该方法显著优于当前最优基线,在高方差随机延迟下仍能实现稳健稳定的遥操作。
原文摘要 · Abstract (English)
Stochastic communication delays in teleoperation introduce signal discontinuities that undermine control stability and degrade control performance. Consequently, the conventional reinforcement learning (RL) methods struggle with the delayed observations due to the delay-induced observations, leading to high-frequency chattering. To address this, we propose a hybrid control framework, delay-resilient RL, integrating a state estimator utilizing Long Short-Term Memory (LSTM) with a residual RL policy, which is resilient to stochastic delays. The LSTM reconstructs smooth, continuous state estimates from delayed observations, enabling the RL agent to learn a residual torque compensation policy that balances tracking accuracy with velocity smoothness. Experimental validation on Franka Panda robots demonstrates that our approach significantly outperforms the state-of-the-art baselines, ensuring robust and stable teleoperation even under high-variance stochastic delays.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。