arXiv:2604.09159cs.LG2026-04

提出一种支持单步采样的多模态强化学习策略,解决传统方法采样慢、训练不稳的问题。

Truncated Rectified Flow Policy for Reinforcement Learning with One-Step Sampling

  • 采用混合确定-随机架构,通过梯度截断和流修正实现稳定训练
  • 在10个MuJoCo任务中多数表现优于基线,单步采样仍具竞争力
  • 适合需要快速推理的复杂多模态决策场景

最大熵强化学习已成为序列决策的标准框架,但其标准高斯策略参数化本质上是单峰的,难以建模复杂的多模态动作分布。这一局限性促使人们越来越多关注基于扩散和流匹配的生成式策略作为更丰富的替代方案。然而,将此类策略融入最大熵强化学习面临两大挑战:连续时间生成策略的似然和熵通常不可计算,且多步采样会引入长时序反向传播不稳定性和显著推理延迟。为解决这些问题,我们提出截断修正流策略(TRFP),该框架基于混合确定-随机架构。这种设计使带熵正则化的优化可计算,同时通过梯度截断和流修正支持稳定训练和有效的单步采样。在玩具多目标环境和10个MuJoCo基准上的实验表明,TRFP能有效捕捉多模态行为,在多数基准上标准采样下优于强基线,且在单步采样下依然具有高度竞争力。

原文摘要 · Abstract (English)

Maximum entropy reinforcement learning (MaxEnt RL) has become a standard framework for sequential decision making, yet its standard Gaussian policy parameterization is inherently unimodal, limiting its ability to model complex multimodal action distributions. This limitation has motivated increasing interest in generative policies based on diffusion and flow matching as more expressive alternatives. However, incorporating such policies into MaxEnt RL is challenging for two main reasons: the likelihood and entropy of continuous-time generative policies are generally intractable, and multi-step sampling introduces both long-horizon backpropagation instability and substantial inference latency. To address these challenges, we propose Truncated Rectified Flow Policy (TRFP), a framework built on a hybrid deterministic-stochastic architecture. This design makes entropy-regularized optimization tractable while supporting stable training and effective one-step sampling through gradient truncation and flow straightening. Empirical results on a toy multigoal environment and 10 MuJoCo benchmarks show that TRFP captures multimodal behavior effectively, outperforms strong baselines on most benchmarks under standard sampling, and remains highly competitive under one-step sampling.

强化学习多模态策略单步采样流匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。