arXiv:2502.12674cs.ROcs.LG2025-02被引 9

受动物运动启发,直接控制关节力矩,提升机器人在复杂环境中的安全与适应性。

SATA: Safe and Adaptive Torque-Based Locomotion Policies Inspired by Animal Learning

  • 模仿动物运动机制,直接输出关节力矩而非角度,增强环境交互能力。
  • 在软地、窄道等挑战场景下零样本迁移成功,无需额外训练。
  • 适合需高安全性的人机共存场景,如服务机器人或救援任务。

尽管基于学习的足式机器人控制器取得进展,但在人机共存环境中部署仍受限于安全问题。现有方法多采用位置控制,策略输出目标关节角,需经低层控制器(如PD或阻抗控制器)转换为力矩。虽在受控真实场景表现优异,但面对训练中未见的环境或扰动时,常因缺乏柔顺性和适应性而产生极端或不安全行为。受动物通过控制肌肉伸缩实现平滑自适应运动的启发,力矩控制策略可直接在力矩空间精确调控执行器,更有效与环境交互,提升安全与适应性。然而,状态空间高度非线性及训练探索效率低等问题阻碍其广泛应用。为此,我们提出SATA,一种受生物力学原理和自适应学习机制启发的框架。该方法显著提升早期探索效率,获得高性能最终策略。实验表明,SATA实现零样本模拟到现实迁移,在软地、滑地、狭窄通道及强外部扰动下均表现出卓越柔顺性与安全性,具备在人机共存及高安全要求场景中实际部署的潜力。

原文摘要 · Abstract (English)

Despite recent advances in learning-based controllers for legged robots, deployments in human-centric environments remain limited by safety concerns. Most of these approaches use position-based control, where policies output target joint angles that must be processed by a low-level controller (e.g., PD or impedance controllers) to compute joint torques. Although impressive results have been achieved in controlled real-world scenarios, these methods often struggle with compliance and adaptability when encountering environments or disturbances unseen during training, potentially resulting in extreme or unsafe behaviors. Inspired by how animals achieve smooth and adaptive movements by controlling muscle extension and contraction, torque-based policies offer a promising alternative by enabling precise and direct control of the actuators in torque space. In principle, this approach facilitates more effective interactions with the environment, resulting in safer and more adaptable behaviors. However, challenges such as a highly nonlinear state space and inefficient exploration during training have hindered their broader adoption. To address these limitations, we propose SATA, a bio-inspired framework that mimics key biomechanical principles and adaptive learning mechanisms observed in animal locomotion. Our approach effectively addresses the inherent challenges of learning torque-based policies by significantly improving early-stage exploration, leading to high-performance final policies. Remarkably, our method achieves zero-shot sim-to-real transfer. Our experimental results indicate that SATA demonstrates remarkable compliance and safety, even in challenging environments such as soft/slippery terrain or narrow passages, and under significant external disturbances, highlighting its potential for practical deployments in human-centric and safety-critical scenarios.

足式机器人力矩控制仿生学习安全控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。