arXiv:2604.22911cs.RO2026-04

一个端到端机器人恢复策略,能自动选择踩步、触墙等动作应对干扰。

RecoverFormer: End-to-End Contact-Aware Recovery for Humanoid Robots

论文配图:RecoverFormer: End-to-End Contact-Aware Recovery for Humanoid Robots
图 1 · 摘自论文原文
  • 用因果Transformer处理50步观测,自适应切换恢复策略。
  • 零样本迁移至有墙环境,100次冲击下成功率100%。
  • 无需人工标注模式,可自动识别不同受力场景的应对方式。

在非结构化环境中运行的人形机器人必须具备应对意外扰动的恢复能力,这一能力对端到端控制策略仍具挑战。本文提出RECOVERFORMER,一种全端到端的人形机器人恢复策略,能够学习在何时以及如何在补偿性跨步、手与环境接触、质心重塑等恢复行为间切换,同时保持在模型失配下的鲁棒性能。该架构结合了对50步观测历史的因果Transformer,以及两个新设计的头部:潜空间恢复模式头,实现不同恢复策略间的平滑过渡;接触可用性头,预测哪些环境表面(墙、扶手、桌边)有助于稳定。我们在MuJoCo中的Unitree G1人形机器人上评估了RECOVERFORMER。仅在开放地面训练,该策略实现零样本迁移至有墙环境,在100–300 N推力及0.25–1.4 m墙距条件下均达100%恢复成功率。在零样本动力学失配下,质量增加25%时成功率75.5%,延迟30 ms下89%,低摩擦下91.5%,复合扰动(摩擦、延迟、质量)下99%。通过t-SNE分析300个任务实例,验证了学习到的潜空间模式在不同力域下的特异性。结果表明,单一端到端策略可实现多模态、接触感知的人形机器人恢复,并泛化于扰动强度、接触几何与动力学偏移。

原文摘要 · Abstract (English)

Humanoid robots operating in unstructured environments must recover from unexpected disturbances-a capability that remains challenging for end-to-end control policies. We present RECOVERFORMER, a fully end-to-end humanoid recovery policy that learns when and how to switch among recovery behaviors-including compensatory stepping, hand-environment contact, and center-of-mass reshaping-while maintaining robust performance under model mismatch. The architecture combines a causal transformer over a 50-step observation history with two novel heads: a latent recovery mode that enables smooth transitions among distinct recovery strategies, and a contact affordance head that predicts which environmental surfaces (walls, railings, table edges) are beneficial for stabilization. We evaluate RECOVERFORMER on the Unitree G1 humanoid in MuJoCo. Trained only on open floor, RECOVERFORMER transfers zero shot to walled environments, achieving 100% recovery success across 100-300 N pushes and across wall distances from 0.25-1.4m. Under zero-shot dynamics mismatch, RECOVERFORMER reaches 75.5% at plus +25% mass, 89% under 30 ms latency, 91.5% at low friction, and 99% under compound friction, latency and mass perturbation. The learned latent modes specialize across force regimes without mode-level supervision, validated by t-SNE analysis of 300 episodes. Taken together, these results show that a single end-to-end policy can deliver multi-modal, contact aware humanoid recovery that generalizes across perturbation magnitude, contact geometry, and dynamics shift.

人形机器人恢复策略端到端接触感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。