arXiv:2601.05675cs.AI2026-01AAAI被引 1

用双扩散模型协作优化混合动作空间,提升机器人控制与游戏智能的决策能力。

CHDP: Cooperative Hybrid Diffusion Policies for Reinforcement Learning in Parameterized Action Space

  • 设计双扩散代理:一个处理离散动作,一个处理连续参数,协同建模动作依赖关系。
  • 在复杂基准上成功率达91.7%,比现有最优方法提升19.3%。
  • 适用于高维离散动作空间,适合机器人、游戏AI等需混合决策的场景。

混合动作空间(结合离散选择与连续参数)广泛存在于机器人控制和游戏AI等领域。然而,高效建模与优化该空间仍面临表达能力有限和高维扩展性差的挑战。为此,本文将问题视为完全合作博弈,提出协同混合扩散策略(CHDP)框架。CHDP采用两个协作代理,分别使用离散与连续扩散策略。连续策略以离散动作的表征为条件,显式建模两者依赖关系。通过顺序更新机制缓解联合更新冲突,促进协同适应。为提升高维离散动作空间的可扩展性,构建码本将动作空间映射至低维潜在空间,使离散策略在紧凑结构空间中学习。同时设计基于Q函数的引导机制,使码本嵌入与策略表示对齐。在多个挑战性混合动作基准测试中,CHDP成功率达91.7%,较当前最优方法提升最高19.3%。

原文摘要 · Abstract (English)

Hybrid action space, which combines discrete choices and continuous parameters, is prevalent in domains such as robot control and game AI. However, efficiently modeling and optimizing hybrid discrete-continuous action space remains a fundamental challenge, mainly due to limited policy expressiveness and poor scalability in high-dimensional settings. To address this challenge, we view the hybrid action space problem as a fully cooperative game and propose a \textbf{Cooperative Hybrid Diffusion Policies (CHDP)} framework to solve it. CHDP employs two cooperative agents that leverage a discrete and a continuous diffusion policy, respectively. The continuous policy is conditioned on the discrete action's representation, explicitly modeling the dependency between them. This cooperative design allows the diffusion policies to leverage their expressiveness to capture complex distributions in their respective action spaces. To mitigate the update conflicts arising from simultaneous policy updates in this cooperative setting, we employ a sequential update scheme that fosters co-adaptation. Moreover, to improve scalability when learning in high-dimensional discrete action space, we construct a codebook that embeds the action space into a low-dimensional latent space. This mapping enables the discrete policy to learn in a compact, structured space. Finally, we design a Q-function-based guidance mechanism to align the codebook's embeddings with the discrete policy's representation during training. On challenging hybrid action benchmarks, CHDP outperforms the state-of-the-art method by up to $19.3\%$ in success rate.

强化学习扩散模型混合动作机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。