arXiv:2606.27766cs.LGcs.AI2026-06

让机器人在安全与风险间灵活权衡,生成更可靠的决策路径。

RS-Diffuser: Risk-Sensitive Diffusion Planning with Distributional Value Guidance

  • 用扩散模型生成多模式轨迹,结合分布价值评估风险。
  • 通过尾部感知目标引导生成,实现不同风险偏好下的行为切换。
  • 在真实场景中显著提升安全性,减少高危动作发生。

离线强化学习允许从固定数据集学习策略而无需额外环境交互,适用于在线探索成本高或危险的安全关键场景。基于扩散的决策方法近年来在离线RL中表现出色,能够建模丰富且多模态的轨迹分布。然而,现有扩散规划器通常为风险中性,可能忽略现实中至关重要的罕见但灾难性结果。本文提出RS-Diffuser,一种融合扩散轨迹生成与分布价值批判的敏感风险离线扩散规划框架。该模型学习未来状态轨迹的扩散规划器、独立的逆动力学模型用于动作解码,以及通过分位数回归估计候选计划完整回报分布的蒙特卡洛分布价值批评器。采样时,将风险敏感引导信号融入去噪过程,利用条件风险价值等尾部感知目标计算梯度,引导生成向期望风险分布偏移。因此,仅需改变推理时的风险参数,单一训练模型即可灵活生成风险规避、风险中性或风险追逐行为。在风险敏感的D4RL和高风险机器人导航基准上的大量实验表明,RS-Diffuser达到当前最佳性能,在提升整体回报的同时增强最坏情况鲁棒性,并减少安全违规。

原文摘要 · Abstract (English)

Offline reinforcement learning enables policy learning from fixed datasets without additional environment interaction, making it appealing for safety-critical applications where online exploration is costly or unsafe. Diffusion-based decision-making methods have recently achieved strong performance in offline RL by modeling rich, multimodal trajectory distributions. However, existing diffusion planners are typically risk-neutral and therefore may overlook rare but catastrophic outcomes that are crucial in real-world deployment. In this work, we propose RS-Diffuser, a risk-sensitive offline diffusion planning framework that combines diffusion-based trajectory generation with distributional value critics. RS-Diffuser learns a diffusion planner over future state trajectories, a separate inverse dynamics model for action decoding, and a Monte Carlo distributional critic that estimates the full return distribution of candidate plans through quantile regression. At sampling time, we incorporate a risk-sensitive guidance signal into the denoising process, using gradients computed from tail-aware objectives such as Conditional Value at Risk to steer generation toward desired risk profiles. As a result, a single trained model can flexibly produce risk-averse, risk-neutral, or risk-seeking behaviors by changing only the inference-time risk parameter. Extensive experiments on risk-sensitive D4RL and risky robot navigation benchmarks demonstrate that RS-Diffuser achieves state-of-the-art performance, improving both overall return and worst-case robustness while reducing safety violations.

扩散模型离线强化学习风险敏感机器人规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。