arXiv:2509.22963cs.LG2025-09被引 8

用扩散模型做组合动作空间的强化学习,更稳更高效。

Reinforcement Learning with Discrete Diffusion Policies for Combinatorial Action Spaces

  • 用镜像下降定义稳定目标分布,让扩散模型匹配它
  • 在多个组合任务上达到领先性能,样本效率显著提升
  • 适合需要复杂决策的大规模强化学习场景

强化学习在大规模组合动作空间中难以扩展,本文提出一种新框架,将离散扩散模型作为高效策略。核心创新是高效的在线训练过程,通过策略镜像下降(PMD)定义正则化的目标策略分布,将策略更新建模为分布匹配问题,使表达能力强的扩散模型能复现这一稳定目标。该解耦方法稳定了学习过程,显著提升了训练性能。在包括DNA序列生成、宏观动作强化学习和多智能体系统在内的多个挑战性组合基准测试中,该方法均取得当前最优结果,且样本效率远超基线模型。

原文摘要 · Abstract (English)

Reinforcement learning (RL) struggles to scale to large, combinatorial action spaces common in many real-world problems. This paper introduces a novel framework for training discrete diffusion models as highly effective policies in these complex settings. Our key innovation is an efficient online training process that ensures stable and effective policy improvement. By leveraging policy mirror descent (PMD) to define an ideal, regularized target policy distribution, we frame the policy update as a distributional matching problem, training the expressive diffusion model to replicate this stable target. This decoupled approach stabilizes learning and significantly enhances training performance. Our method achieves state-of-the-art results and superior sample efficiency across a diverse set of challenging combinatorial benchmarks, including DNA sequence generation, RL with macro-actions, and multi-agent systems. Experiments demonstrate that our diffusion policies attain superior performance compared to other baselines.

强化学习扩散模型组合优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。