将流匹配拓展到离线强化学习的离散动作空间,支持多目标决策。
Flow Matching for Offline Reinforcement Learning with Discrete Actions
- 用连续时间马尔可夫链替代连续流,通过Q加权目标训练。
- 在高维控制、多智能体博弈等场景中性能超越传统方法。
- 适用于离散与量化连续动作,适合多模态决策场景。
基于扩散模型和流匹配的生成策略在离线强化学习中展现出强大潜力,但其应用仍主要局限于连续动作空间。为扩展至更广泛的离线强化学习场景,我们提出一种通用框架,将流匹配推广至支持离散动作空间的多目标设置。具体地,用连续时间马尔可夫链替代连续流,并采用Q加权流匹配目标进行训练。进一步地,该设计被拓展至多智能体场景,通过因子化条件路径缓解联合动作空间的指数增长问题。理论上,在理想条件下优化该目标可恢复最优策略。大量实验表明,该方法在多样化任务和基准测试中表现稳健,涵盖高维控制、多智能体博弈及动态偏好变化的多目标场景,且在实际多模态决策中优于传统离线强化学习方法。此外,通过动作量化,该离散框架亦可应用于连续控制问题,实现表示复杂度与性能间的灵活权衡。
原文摘要 · Abstract (English)
Generative policies based on diffusion models and flow matching have shown strong promise for offline reinforcement learning (RL), but their applicability remains largely confined to continuous action spaces. To address a broader range of offline RL settings, we extend flow matching to a general framework that supports discrete action spaces with multiple objectives. Specifically, we replace continuous flows with continuous-time Markov chains, trained using a Q-weighted flow matching objective. We then extend our design to multi-agent settings, mitigating the exponential growth of joint action spaces via a factorized conditional path. We theoretically show that, under idealized conditions, optimizing this objective recovers the optimal policy. Extensive experiments further demonstrate that our method performs robustly across diverse settings and benchmarks, including high-dimensional control, multi-agent games, and dynamically changing preferences over multiple objectives, while outperforming traditional offline RL methods in practical multi-modal decision-making scenarios. Our discrete framework can also be applied to continuous-control problems through action quantization, providing a flexible trade-off between representational complexity and performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。