让仿人机器人在复杂地形上跌倒后自动恢复,无需额外训练
VIGOR: Visual Goal-In-Context Inference for Unified Humanoid Fall Safety
- 用人类跌倒动作的共性规律指导机器人学习,实现跨地形通用恢复
- 在真实Unitree G1机器人上实现零样本跌倒安全,无需实测微调
- 结合视觉与身体感知,让机器人快速做出整体反应,适合复杂环境任务
在杂乱环境中,可靠的跌倒恢复对仿人机器人至关重要。与四足或轮式机器人不同,仿人机器人在跌倒过程中面临高能冲击、全身接触和大幅视角变化,恢复能力直接影响持续作业。现有方法将跌倒安全拆分为避障、减震和站起等独立问题,或依赖无视觉的端到端策略,通常仅在平坦地形上训练。深层来看,跌倒安全被视作单一数据复杂性,耦合姿态、动力学与地形信息,导致泛化性差。本文提出统一的跌倒安全框架,涵盖跌倒全过程。基于两个洞察:1)人类跌倒与恢复姿态高度受限且可迁移,通过姿态对齐即可从平坦地形推广至复杂地形;2)快速全身反应需整合感知与运动表征。我们使用稀疏人类示范在平坦与模拟复杂地形上训练一个特权教师模型,并将其知识蒸馏到仅依赖本体感觉与前景深度的可部署学生模型。学生通过匹配教师的“目标-上下文”隐空间表示(融合目标姿态与局部地形),而非分别编码感知与行为。仿真与真实单位元G1机器人实验表明,该方法可在多样化非平坦环境中实现鲁棒的零样本跌倒安全,无需实测微调。
原文摘要 · Abstract (English)
Reliable fall recovery is critical for humanoids operating in cluttered environments. Unlike quadrupeds or wheeled robots, humanoids experience high-energy impacts, complex whole-body contact, and large viewpoint changes during a fall, making recovery essential for continued operation. Existing methods fragment fall safety into separate problems such as fall avoidance, impact mitigation, and stand-up recovery, or rely on end-to-end policies trained without vision through reinforcement learning or imitation learning, often on flat terrain. At a deeper level, fall safety is treated as monolithic data complexity, coupling pose, dynamics, and terrain and requiring exhaustive coverage, limiting scalability and generalization. We present a unified fall safety approach that spans all phases of fall recovery. It builds on two insights: 1) Natural human fall and recovery poses are highly constrained and transferable from flat to complex terrain through alignment, and 2) Fast whole-body reactions require integrated perceptual-motor representations. We train a privileged teacher using sparse human demonstrations on flat terrain and simulated complex terrains, and distill it into a deployable student that relies only on egocentric depth and proprioception. The student learns how to react by matching the teacher's goal-in-context latent representation, which combines the next target pose with the local terrain, rather than separately encoding what it must perceive and how it must act. Results in simulation and on a real Unitree G1 humanoid demonstrate robust, zero-shot fall safety across diverse non-flat environments without real-world fine-tuning. The project page is available at https://vigor2026.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。