用价值函数引导流模型,实现高维连续控制的高效探索。
Scalable Exploration for High-Dimensional Continuous Control via Value-Guided Flow
- 基于价值函数构建概率流,沿任务相关方向探索动作空间。
- 在多个高维控制任务上超越主流强化学习基线,样本效率显著提升。
- 适用于全身体感肌肉模型,支持复杂敏捷动作的实时控制。
生物与机器人系统中的高维连续控制因状态-动作空间庞大而面临挑战,有效探索至关重要。传统强化学习探索策略多为无向性,随动作维度增加性能急剧下降。许多方法依赖降维,牺牲策略表达力与系统灵活性。本文提出Q引导流探索(Qflex),一种直接在原始高维动作空间中进行探索的可扩展强化学习方法。训练过程中,Qflex从可学习的源分布出发,沿由学习到的价值函数诱导的概率流遍历动作,使探索方向与任务相关梯度对齐,而非依赖各向同性噪声。该方法在多个高维连续控制基准任务上显著优于代表性在线强化学习基线。Qflex成功控制了全身体感肌肉模型,完成复杂敏捷运动,展现出在极高维场景下的优越可扩展性与样本效率。结果表明,价值引导流为大规模探索提供了原理清晰且实用的路径。
原文摘要 · Abstract (English)
Controlling high-dimensional systems in biological and robotic applications is challenging due to expansive state-action spaces, where effective exploration is critical. Commonly used exploration strategies in reinforcement learning are largely undirected with sharp degradation as action dimensionality grows. Many existing methods resort to dimensionality reduction, which constrains policy expressiveness and forfeits system flexibility. We introduce Q-guided Flow Exploration (Qflex), a scalable reinforcement learning method that conducts exploration directly in the native high-dimensional action space. During training, Qflex traverses actions from a learnable source distribution along a probability flow induced by the learned value function, aligning exploration with task-relevant gradients rather than isotropic noise. Our proposed method substantially outperforms representative online reinforcement learning baselines across diverse high-dimensional continuous-control benchmarks. Qflex also successfully controls a full-body human musculoskeletal model to perform agile, complex movements, demonstrating superior scalability and sample efficiency in very high-dimensional settings. Our results indicate that value-guided flows offer a principled and practical route to exploration at scale.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。