对比多种强化学习算法在传感器失效环境下的导航表现,发现DreamerV3更稳健。
Benchmarking Deep Reinforcement Learning for Navigation in Denied Sensor Environments
- 构建可配置传感器干扰的导航基准测试,评估不同算法鲁棒性。
- DreamerV3在动态目标视觉导航中表现最优,其他方法无法学习该任务。
- 对抗训练提升传感器缺失场景表现,但牺牲常规环境性能。
深度强化学习(DRL)被用于未知环境中的自主导航。现有研究多假设传感器数据完美,但真实环境常存在自然与人为噪声及信号中断。本文构建了一个可配置传感器干扰的导航基准,评估主流与新兴DRL算法的表现。重点比较了无模型PPO与基于模型的DreamerV3在传感器受限条件下的差异。结果表明,DreamerV3在动态目标视觉端到端导航任务中显著优于其他方法,且其他算法无法完成该任务;在各类传感器缺失场景下,DreamerV3整体表现更优。为提升鲁棒性,引入对抗训练,虽在传感器正常环境上性能略有下降,但在干扰环境中效果显著改善。本工作为开发应对不确定与传感器失效的先进导航策略提供了起点。
原文摘要 · Abstract (English)
Deep Reinforcement learning (DRL) is used to enable autonomous navigation in unknown environments. Most research assume perfect sensor data, but real-world environments may contain natural and artificial sensor noise and denial. Here, we present a benchmark of both well-used and emerging DRL algorithms in a navigation task with configurable sensor denial effects. In particular, we are interested in comparing how different DRL methods (e.g. model-free PPO vs. model-based DreamerV3) are affected by sensor denial. We show that DreamerV3 outperforms other methods in the visual end-to-end navigation task with a dynamic goal - and other methods are not able to learn this. Furthermore, DreamerV3 generally outperforms other methods in sensor-denied environments. In order to improve robustness, we use adversarial training and demonstrate an improved performance in denied environments, although this generally comes with a performance cost on the vanilla environments. We anticipate this benchmark of different DRL methods and the usage of adversarial training to be a starting point for the development of more elaborate navigation strategies that are capable of dealing with uncertain and denied sensor readings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。