通过扩展特征与聚类方法,揭示了运动策略中的隐含周期性阶段结构。
Visualizing Latent Phase Structures in Locomotion Policies: A Multi-Environment Study with Temporal Feature Extension

- 融合状态、动作、下一状态和下一动作构建增强特征
- 在三个环境上识别出更清晰的运动相位过渡规则
- 适合对强化学习策略可解释性感兴趣的研究者
深度强化学习(DRL)在MuJoCo基准任务如HalfCheetah、Ant和Walker2D上表现出色。然而,如何可视化由深度神经网络实现的训练策略内部获得的运动结构仍具挑战。生物力学研究表明,运动控制依赖于支撑相与摆动相等重复性运动阶段。本研究提出一种框架,通过与环境交互生成的轨迹,揭示运动控制策略中的潜在运动相结构。该方法将聚类特征从仅状态观测扩展至包含动作、下一状态和下一动作的增强特征,并提出一种抑制自转移的聚类数量确定方法。在Ant-v5、HalfCheetah-v5和Walker2D-v5三个环境中应用后,成功识别出比现有方法更清晰、更规则的相位转换模式。
原文摘要 · Abstract (English)
Deep reinforcement learning (DRL) has been shown to achieve high performance on locomotion control tasks in MuJoCo benchmarks such as HalfCheetah, Ant, and Walker2D. However, visualizing the motion structures internally obtained by a trained policy function implemented as a deep neural network remains challenging. It is known from biomechanics and related fields that locomotion control is realized through the repetition of motion phases such as the stance phase and swing phase. In this study, we propose a framework for uncovering latent motion phase structures from trajectories generated by locomotion control policies through interaction with the environment. The proposed method extends the clustering features from state observations alone to augmented features including actions, next states, and next actions, and introduces a method for determining the number of clusters that suppresses self-transitions. Applying the proposed method to three environments -- Ant-v5, HalfCheetah-v5, and Walker2D-v5 -- we successfully identified phase structures with clearer and more regular transition rules than those obtained by the existing method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。