arXiv:2606.12640cs.LGcs.RO2026-06中稿 · the 23rd IFAC Worl…

用神经控制屏障函数让扩散模型生成更安全的多智能体策略。

Individual Control Barrier Functions-Guided Diffusion Model for Safe Offline Multi-Agent Reinforcement Learning

论文配图:Individual Control Barrier Functions-Guided Diffusion Model for Safe Offline Multi-Agent Reinforcement Learning
图 1 · 摘自论文原文
  • 将神经控制屏障函数嵌入扩散模型,引导安全轨迹生成
  • 在多个基准上实现显著安全提升,同时保持高奖励水平
  • 适合需要安全保障的多智能体系统研究者

离线强化学习允许直接从数据中学习控制策略而无需在线交互,适用于安全关键任务。近期研究将扩散模型应用于离线强化学习,利用其强大的复杂数据分布建模能力。然而,现有方法主要聚焦于单智能体场景,多智能体环境中的安全挑战仍缺乏探索。本文提出一种安全的离线多智能体强化学习算法,通过将神经个体控制屏障函数嵌入扩散模型,在轨迹生成过程中增强安全性,并通过逆动力学恢复控制策略。我们在多种基准上评估该算法,结果表明其在保持竞争力奖励的同时实现了显著的安全性提升。

原文摘要 · Abstract (English)

Offline reinforcement learning allows control policies to be learned directly from data without online interaction, making it suitable for safety-critical tasks. Recent studies have applied diffusion models to offline reinforcement learning to leverage their strong capacity for modeling complex data distributions. However, existing approaches primarily focus on single-agent settings, leaving the safety challenges in multi-agent environments largely unexplored. In this work, we propose a safe offline multi-agent reinforcement learning algorithm that embeds neural individual control barrier functions into the diffusion model to enhance safety during trajectory generation, with control policies recovered through inverse dynamics. We evaluate our algorithm across diverse benchmarks, demonstrating substantial safety improvements while maintaining competitive rewards.

多智能体扩散模型安全强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。