让机器人提前预判轨迹,减少控制延迟和超调。
Anticipatory Reinforcement Learning for Trajectory Tracking
- 在强化学习状态中加入目标速度和未来参考点,实现前瞻控制。
- 仿真中误差降低9倍,均值绝对偏差从2.73°降至0.31°。
- 简单配置即可实现实体硬件最优表现,适合工业控制部署。
工业控制中的深度强化学习常因仅基于当前跟踪误差的反应式控制而出现滞后与超调。为实现低计算开销的前瞻控制,我们提出一种预测性方法,将目标速度和未来参考时域引入DRL状态空间。在单自由度直升机测试平台使用近端策略优化(PPO)评估八种配置,仿真结果显示误差降低9倍,均值绝对偏差从2.73°降至0.31°。然而,零样本迁移至物理硬件时暴露出仿真到现实的差距。有趣的是,仅使用单一远距离前瞻时域的简化配置,其在真实场景中的表现与最复杂模型相当(1.11°)。整体表明,高粒度预测数据并非物理迁移的必要条件。
原文摘要 · Abstract (English)
Deep reinforcement learning (DRL) in industrial control often suffers from lag and overshoot due to purely reactive control based on the current tracking error. To achieve anticipatory control without high computational overhead, we introduce a predictive formulation that augments the DRL state space with target velocities and future reference horizons. Evaluating eight configurations using proximal policy optimization (PPO) on a 1-degree-of-freedom (1-DoF) helicopter testbed, simulation results showed a 9-fold error reduction, lowering the mean absolute deviation from 2.73° to 0.31°. However, zero-shot transfer to physical hardware revealed a sim-to-real gap. Interestingly, a simpler configuration using a single, further look-ahead horizon matched the real-world top performance of the most complex model (1.11°). Overall, evaluating various combinations of prediction horizons and target velocities demonstrated that highly granular predictive data is not necessarily required for physical transfer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。