arXiv:2507.12977cs.RO2025-07中稿 · IROS 2025被引 3

让扩散模型直接优化安全与效率,提升机器人避障与到达能力。

Non-differentiable Reward Optimization for Diffusion-based Autonomous Motion Planning

  • 用强化学习+奖励加权动态阈值,训练扩散模型直接优化不可导目标。
  • 在CrowdNav和ETH-UCY数据集上优于传统可微方法,提升避障与达目标率。
  • 适合需要高安全性的自动驾驶、人机交互场景的运动规划研究者。

安全有效的运动规划对自主机器人至关重要。扩散模型擅长捕捉复杂智能体交互,是动态环境中决策的核心。近期研究已成功将扩散模型应用于运动规划,展现出处理复杂场景及准确预测多模态未来轨迹的能力。然而,扩散模型在训练目标上存在局限,其仅近似数据分布,未显式建模底层决策动态。而运动规划的关键在于不可导的下游目标,如安全(碰撞避免)与有效性(目标到达),传统学习算法无法直接优化这些目标。本文提出一种基于强化学习的扩散运动规划模型训练方案,使其能有效学习显式衡量安全与有效性的非可导目标。具体而言,引入奖励加权动态阈值算法,构建密集奖励信号,显著提升训练效果,优于基于可微目标的模型。在行人数据集CrowdNav和ETH-UCY上的实验表明,该方法在安全与有效性方面均达到最先进性能,验证了其在复杂场景中规划的通用性。

原文摘要 · Abstract (English)

Safe and effective motion planning is crucial for autonomous robots. Diffusion models excel at capturing complex agent interactions, a fundamental aspect of decision-making in dynamic environments. Recent studies have successfully applied diffusion models to motion planning, demonstrating their competence in handling complex scenarios and accurately predicting multi-modal future trajectories. Despite their effectiveness, diffusion models have limitations in training objectives, as they approximate data distributions rather than explicitly capturing the underlying decision-making dynamics. However, the crux of motion planning lies in non-differentiable downstream objectives, such as safety (collision avoidance) and effectiveness (goal-reaching), which conventional learning algorithms cannot directly optimize. In this paper, we propose a reinforcement learning-based training scheme for diffusion motion planning models, enabling them to effectively learn non-differentiable objectives that explicitly measure safety and effectiveness. Specifically, we introduce a reward-weighted dynamic thresholding algorithm to shape a dense reward signal, facilitating more effective training and outperforming models trained with differentiable objectives. State-of-the-art performance on pedestrian datasets (CrowdNav, ETH-UCY) compared to various baselines demonstrates the versatility of our approach for safe and effective motion planning.

扩散模型运动规划强化学习机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。