arXiv:2601.22211cs.LG2026-01被引 1

用球面流生成可满足约束的组合动作,提升强化学习表现。

Latent Spherical Flow Policy for Reinforcement Learning with Combinatorial Actions

  • 在紧凑连续潜空间用球面流生成随机策略,由求解器保证动作可行性。
  • 在多个任务上平均性能优于基线20.6%,且训练更高效。
  • 适合需要复杂约束下决策的强化学习场景,如资源分配、路径规划。

带有组合动作空间的强化学习仍具挑战性,因可行动作集呈指数级增长且受复杂可行性约束,直接参数化策略不可行。现有方法将特定任务的价值函数嵌入约束优化问题或学习确定性结构化策略,牺牲了通用性与策略表达能力。本文提出一种由求解器引导的潜在球面流策略(LSFlow),在保持现代生成策略表达力的同时,通过设计保证可行性。该方法在紧凑连续潜空间中通过球面流匹配学习随机策略,并将每个潜变量样本映射到有效结构化动作,交由组合优化求解器完成。为提高效率,价值网络在潜空间直接训练,避免策略优化时重复调用求解器。针对求解器导致的价值函数分段常数和不连续问题,引入平滑贝尔曼算子以获得稳定学习目标。实验表明,该方法在多个具有挑战性的组合强化学习任务上平均性能超越当前最优基线20.6%。

原文摘要 · Abstract (English)

Reinforcement learning (RL) with combinatorial action spaces remains challenging because feasible action sets are exponentially large and governed by complex feasibility constraints, making direct policy parameterization impractical. Existing approaches embed task-specific value functions into constrained optimization programs or learn deterministic structured policies, sacrificing generality and policy expressiveness. We propose a solver-induced \emph{latent spherical flow policy} that brings the expressiveness of modern generative policies to combinatorial RL while guaranteeing feasibility by design. Our method, LSFlow, learns a \emph{stochastic} policy in a compact continuous latent space via spherical flow matching, and delegates feasibility to a combinatorial optimization solver that maps each latent sample to a valid structured action. To improve efficiency, we train the value network directly in the latent space, avoiding repeated solver calls during policy optimization. To address the piecewise-constant and discontinuous value landscape induced by solver-based action selection, we introduce a smoothed Bellman operator that yields stable, well-defined learning targets. Empirically, our approach outperforms state-of-the-art baselines by an average of 20.6\% across a range of challenging combinatorial RL tasks.

强化学习组合优化生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。