arXiv:2509.23846cs.LGcs.AI2025-09NeurIPS被引 5

用扩散模型生成最坏情况轨迹,提升强化学习鲁棒性。

Adversarial Diffusion for Robust Reinforcement Learning

  • 用条件扩散模型生成最坏情况轨迹,模拟环境不确定性。
  • 在标准基准上优于现有鲁棒强化学习方法,性能更优。
  • 适合对环境扰动敏感的机器人控制等应用场景。

强化学习中的建模误差与不确定性仍是核心挑战。本文利用扩散模型训练鲁棒的强化学习策略:扩散模型能一次性生成完整轨迹,避免逐步转移模型的累积误差;且可通过条件采样从特定分布采样,灵活性高。我们基于条件风险价值(CVaR)优化与鲁棒强化学习的关联,提出对抗性扩散鲁棒强化学习(AD-RRL),在训练中引导扩散过程生成最坏情况轨迹,从而优化累计回报的CVaR。在多个标准基准上的实验表明,相较于现有鲁棒强化学习方法,AD-RRL展现出更优的鲁棒性与性能。

原文摘要 · Abstract (English)

Robustness to modeling errors and uncertainties remains a central challenge in reinforcement learning (RL). In this work, we address this challenge by leveraging diffusion models to train robust RL policies. Diffusion models have recently gained popularity in model-based RL due to their ability to generate full trajectories "all at once", mitigating the compounding errors typical of step-by-step transition models. Moreover, they can be conditioned to sample from specific distributions, making them highly flexible. We leverage conditional sampling to learn policies that are robust to uncertainty in environment dynamics. Building on the established connection between Conditional Value at Risk (CVaR) optimization and robust RL, we introduce Adversarial Diffusion for Robust Reinforcement Learning (AD-RRL). AD-RRL guides the diffusion process to generate worst-case trajectories during training, effectively optimizing the CVaR of the cumulative return. Empirical results across standard benchmarks show that AD-RRL achieves superior robustness and performance compared to existing robust RL methods.

强化学习扩散模型鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。