针对软体机器人在分布外场景下的控制难题,提出抗分布偏移的离线强化学习方法。
DiSA-IQL: Offline Reinforcement Learning for Robust Soft Robot Control under Distribution Shifts
- 通过惩罚不可靠状态动作对,增强离线强化学习对分布偏移的鲁棒性。
- 在模拟环境中实现更高成功率与更平滑轨迹,优于行为克隆等基线方法。
- 适合需要安全训练的柔性机器人控制任务,尤其适用于现实部署场景。
软体蛇形机器人在复杂环境中表现出卓越的灵活性与适应性,但其控制因高度非线性动力学而困难。现有基于模型和仿生的控制器依赖简化假设,性能受限。深度强化学习(DRL)虽具潜力,但在线训练成本高且可能损伤设备。离线强化学习利用预收集数据更安全,却面临分布偏移问题,导致泛化能力下降。为此,我们提出DiSA-IQL(分布偏移感知隐式Q学习),在IQL基础上引入鲁棒性调节机制,通过惩罚不可靠状态-动作对来缓解分布偏移。我们在两个设定下评估了该方法:同分布与分布外测试。仿真结果表明,DiSA-IQL始终优于行为克隆(BC)、保守Q学习(CQL)及原始IQL,在目标达成任务中取得更高成功率、更平滑轨迹和更强鲁棒性。
原文摘要 · Abstract (English)
Soft snake robots offer remarkable flexibility and adaptability in complex environments, yet their control remains challenging due to highly nonlinear dynamics. Existing model-based and bio-inspired controllers rely on simplified assumptions that limit their performance. Deep reinforcement learning (DRL) has recently emerged as a promising alternative, but online training is often impractical because of costly and potentially damaging real-world interactions. Offline RL provides a safer option by leveraging pre-collected datasets, but it suffers from distribution shift, which degrades generalization to unseen scenarios. To overcome this challenge, we propose DiSA-IQL (Distribution-Shift-Aware Implicit Q-Learning), an extension of IQL that incorporates robustness modulation by penalizing unreliable state-action pairs to mitigate distribution shift. We evaluate DiSA-IQL on goal-reaching tasks across two settings: in-distribution and out-of-distribution evaluation. Simulation results show that DiSA-IQL consistently outperforms baseline models, including Behavior Cloning (BC), Conservative Q-Learning (CQL), and vanilla IQL, achieving higher success rates, smoother trajectories, and greater robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。