让机器人在复杂地形上稳走,还能扛15公斤重物不倒。
TACT-ful: Multi-Channel Terrain Affordance and Compliance Training for Payload-Robust Perceptive Humanoid Locomotion

- 用多通道地形信号指导脚点选择,结合姿态与力矩感知。
- 模拟中实现1米/秒行进速度,能应对0.2米高台阶和15公斤负载。
- 无需额外训练,直接从仿真部署到真实人形机器人。
在结构化地形上进行足点选择需显式考虑接触平面性、表面坡度与运动可达性,单一高度信号无法捕捉这些特性。本文提出一种融合平面度、坡度与速度感知高度可行性的多通道地形代价,并引入前向攀爬奖励,同时驱动基于GPU并行的发散分量运动(DCM)足点规划器,塑造密集的每步可操作性奖励,用于训练非对称演员-评论家策略,采用近端策略优化(PPO)从深度图像学习。通过贝塞尔摆动轨迹配合自适应顶点偏置,扩展足点追踪至关节位置与姿态,利用弧切线引导足底在踏板跨越与踏面着陆时的姿态调整。为支持负载任务,引入下肢柔顺性训练:在采样载荷附着点注入虚拟力矩,生成物理一致的力与力矩;以力矩感知的柔顺目标替代刚性姿态惩罚,使策略学会在无力感测条件下响应负载扰动。整个系统使用标准PPO端到端训练,无需知识蒸馏或师生阶段,仅通过配置变更即可直接部署于真实人形机器人。仿真结果表明,该策略在楼梯上达到1.0米/秒速度,可处理最高0.20米的台阶,且负载鲁棒性提升至约15公斤中心负载及力矩主导的手腕负载,无需微调。同时提供结构化地形上的定性硬件演示。
原文摘要 · Abstract (English)
Foothold selection on structured terrain requires explicit reasoning about contact planarity, surface steepness, and kinematic reachability, properties not captured by a single height-based terrain signal. We propose a multi-channel terrain cost combining flatness, steepness, and velocity-aware height feasibility, plus a forward climb reward, that simultaneously drives a GPU-parallel divergent component of motion (DCM) foothold planner and shapes a dense per-step affordance reward for an asymmetric actor-critic policy trained with proximal policy optimization (PPO) from depth images. A Bézier swing trajectory with adaptive apex bias extends foothold tracking to joint position-and-orientation, using the arc tangent to guide sole orientation through riser crossings and tread landings. To support payload tasks, we introduce a lower-body compliance training procedure in which a virtual wrench is injected at a sampled load attachment point, generating physically consistent force and moment; wrench-aware compliance targets replace rigid pose penalties, and the policy learns to yield to load-induced perturbations without force sensing. The full system trains end-to-end with standard PPO, no distillation, and no teacher-student staging, and is deployed on a humanoid directly from simulation with configuration changes only. In simulation, the policy reaches $1.0~\mathrm{m/s}$ on stairs with risers up to $0.20~\mathrm{m}$ and improves payload robustness up to ${\sim}15~\mathrm{kg}$ centered load and for moment-dominated wrist loads without fine-tuning. We also provide a qualitative hardware demonstration on structured terrain. Project website: https://fai-rl-tech.github.io/tact-locomotion.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。