arXiv:2510.10851cs.RO2025-10

让机器人行走时既能精准跟命令,又能柔性应对外力。

Preference-Conditioned Multi-Objective RL for Integrated Command Tracking and Force Compliance in Humanoid Locomotion

  • 用偏好输入控制策略,在跟命令和抗外力间动态权衡。
  • 仿真与实机测试均稳定,可部署于真实人形机器人。
  • 适合需要人机交互的机器人运动控制场景。

人形机器人行走需兼顾导航指令跟踪与外部作用力下的柔顺响应。现有强化学习方法多侧重鲁棒性,导致机器人抗拒外力但缺乏柔顺性,尤其对本就不稳定的双足机器人构成挑战。本文将行走任务建模为多目标优化问题,平衡指令跟踪与外力柔顺性。提出偏好条件化的多目标强化学习框架,通过用户指定的偏好输入,使单一全向行走策略在两种目标间灵活切换。外力通过速度阻力系数建模以保证奖励设计一致性,训练采用编码器-解码器结构,从可获取观测中推断任务相关特权特征。在仿真与真实人形机器人上验证,结果表明该框架训练稳定,支持可部署的偏好可控行走能力。

原文摘要 · Abstract (English)

Humanoid locomotion requires not only accurate command tracking for navigation but also compliant responses to external forces during human interaction. Despite significant progress, existing RL approaches mainly emphasize robustness, yielding policies that resist external forces but lack compliance particularly challenging for inherently unstable humanoids. In this work, we address this by formulating humanoid locomotion as a multi-objective optimization problem that balances command tracking and external force compliance. We introduce a preference-conditioned multi-objective RL (MORL) framework that enables a single omnidirectional locomotion policy to trade off between command following and force compliance via a user-specified preference input. External forces are modeled via velocity-resistance factor for consistent reward design, and training leverages an encoder-decoder structure that infers task-relevant privileged features from deployable observations. We validate our approach in both simulation and real-world experiments on a humanoid robot. Experimental results in simulation and on hardware show that the framework trains stably and enables deployable preference-conditioned humanoid locomotion.

人形机器人多目标强化学习柔顺控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。