用分层强化学习让外骨骼自适应走路,提升行动能力。
Hierarchical Reinforcement Learning Framework for Adaptive Walking Control Using General Value Functions of Lower-Limb Sensor Signals
- 分层架构:高层决策地形策略,底层用通用价值函数预测传感器信号。
- 加入预测后,系统在平地、斜坡、转弯等场景分类准确率显著提升。
- 适合康复机器人、外骨骼控制研究者,尤其关注智能决策的场景。
康复技术是研究人机协同学习与决策的天然场景。本文探索使用分层强化学习(HRL)开发下肢外骨骼的自适应控制策略,以提升运动障碍者的移动能力与自主性。受生物体感觉运动处理模型启发,所提出的HRL方法将复杂控制任务分解为高层地形策略适应与低层预测信息生成两部分;后者通过持续学习通用价值函数(GVFs)实现。GVFs从多源可穿戴传感器(如肌电、压力鞋垫、角度计)中提取未来信号的时序抽象。我们对比了将实际与预测信号输入策略网络的两种方法,旨在增强外骨骼在不同地形行走时的决策能力。关键结果表明,引入GVF预测显著提升了整体网络精度。在平坦地面、不平路面、上下坡及转弯等场景中均观察到性能提升,这些地形在无预测信息时易被误判。这说明预测信息有助于应对不确定性,例如高误判风险的地形。本工作为理解HRL的细节以及外骨骼安全穿越多样环境的未来发展提供了新见解。
原文摘要 · Abstract (English)
Rehabilitation technology is a natural setting to study the shared learning and decision-making of human and machine agents. In this work, we explore the use of Hierarchical Reinforcement Learning (HRL) to develop adaptive control strategies for lower-limb exoskeletons, aiming to enhance mobility and autonomy for individuals with motor impairments. Inspired by prominent models of biological sensorimotor processing, our investigated HRL approach breaks down the complex task of exoskeleton control adaptation into a higher-level framework for terrain strategy adaptation and a lower-level framework for providing predictive information; this latter element is implemented via the continual learning of general value functions (GVFs). GVFs generated temporal abstractions of future signal values from multiple wearable lower-limb sensors, including electromyography, pressure insoles, and goniometers. We investigated two methods for incorporating actual and predicted sensor signals into a policy network with the intent to improve the decision-making capacity of the control system of a lower-limb exoskeleton during ambulation across varied terrains. As a key result, we found that the addition of predictions made from GVFs increased overall network accuracy. Terrain-specific performance increases were seen while walking on even ground, uneven ground, up and down ramps, and turns, terrains that are often misclassified without predictive information. This suggests that predictive information can aid decision-making during uncertainty, e.g., on terrains that have a high chance of being misclassified. This work, therefore, contributes new insights into the nuances of HRL and the future development of exoskeletons to facilitate safe transitioning and traversing across different walking environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。