自模仿扩散策略提升视觉导航效率与鲁棒性
Self-Imitated Diffusion Policy for Efficient and Robust Visual Navigation
- 通过自采样轨迹选择性模仿,减少对专家数据依赖
- 实测推理速度达110ms,比基线快2.5倍,支持实时部署
- 适用于多机器人平台,尤其适合算力受限场景
扩散策略(DP)在视觉导航中展现出捕捉多模态轨迹分布的潜力。然而,主流DP方法依赖的模仿学习常继承专家示范中的次优与冗余问题,导致推理时需采用计算密集型的‘生成-过滤’流程,依赖额外筛选器。为此,我们提出自模仿扩散策略(SIDP),通过从自身采样轨迹中选择性模仿,实现高效规划。具体地,SIDP引入奖励驱动的自模仿机制,鼓励策略持续输出高质量轨迹,而非质量不一的结果,从而降低对大规模采样和后处理的依赖。训练中采用奖励驱动的课程学习以提升数据利用效率,并通过无目标探索增强轨迹多样性,提升规划鲁棒性。在综合仿真基准上的大量实验表明,SIDP显著优于已有方法;真实机器人实验也验证了其在多个平台上的有效性。在Jetson Orin Nano上,SIDP推理时间仅110ms,相比基线NavDP的273ms提速2.5倍,支持高效实时部署。
原文摘要 · Abstract (English)
Diffusion policies (DP) have demonstrated significant potential in visual navigation by capturing diverse multi-modal trajectory distributions. However, standard imitation learning (IL), which most DP methods rely on for training, often inherits sub-optimality and redundancy from expert demonstrations, thereby necessitating a computationally intensive "generate-then-filter" pipeline that relies on auxiliary selectors during inference. To address these challenges, we propose Self-Imitated Diffusion Policy (SIDP), a novel framework that learns improved planning by selectively imitating a set of trajectories sampled from itself. Specifically, SIDP introduces a reward-guided self-imitation mechanism that encourages the policy to consistently produce high-quality trajectories efficiently, rather than outputs of inconsistent quality, thereby reducing reliance on extensive sampling and post-filtering. During training, we employ a reward-driven curriculum learning paradigm to mitigate inefficient data utility, and goal-agnostic exploration for trajectory augmentation to improve planning robustness. Extensive evaluations on a comprehensive simulation benchmark show that SIDP significantly outperforms previous methods, with real-world experiments confirming its effectiveness across multiple robotic platforms. On Jetson Orin Nano, SIDP delivers a 2.5$\times$ faster inference than the baseline NavDP, i.e., 110ms VS 273ms, enabling efficient real-time deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。