让机器人从视觉感知自主完成复杂动作,无需预先设定轨迹。
ULTRA: Unified Multimodal Control for Autonomous Humanoid Whole-Body Loco-Manipulation
- 用物理驱动算法将动捕数据转为类人机器人可用的自然动作。
- 在真实机器人上实现仅凭视觉输入就完成目标导向的全身协同操作。
- 支持从精准到噪声较大的感知输入,适合实际应用部署。
实现自主且多功能的全身运动与操作仍是使类人机器人实用化的关键挑战。现有方法存在根本局限:动捕数据难以获取或质量低;难以扩展至大规模技能库;更重要的是,依赖预设运动参考进行跟踪,而非基于感知和高层任务指令生成行为。为此,我们提出ULTRA统一框架,包含两个核心组件:首先,提出一种物理驱动的神经重定向算法,将大规模动捕数据转换为类人机器人动作,保持接触密集交互的物理合理性;其次,学习一个统一的多模态控制器,支持密集参考与稀疏任务指令,在从高精度动捕状态到含噪自指视觉输入的多种感知条件下运行。通过将通用追踪策略提炼进控制器,将运动技能压缩至紧凑隐空间,并使用强化学习微调以扩展覆盖范围并提升分布外场景下的鲁棒性。这使得系统可在测试时无需参考轨迹,仅凭稀疏意图即可协调执行全身行为。我们在仿真环境和真实Unitree G1机器人上评估了ULTRA,结果表明其能从自指视觉输入实现自主、目标导向的全身运动与操作,显著优于仅依赖追踪的基线方法,且技能有限。
原文摘要 · Abstract (English)
Achieving autonomous and versatile whole-body loco-manipulation remains a central barrier to making humanoids practically useful. Yet existing approaches are fundamentally constrained: retargeted data are often scarce or low-quality; methods struggle to scale to large skill repertoires; and, most importantly, they rely on tracking predefined motion references rather than generating behavior from perception and high-level task specifications. To address these limitations, we propose ULTRA, a unified framework with two key components. First, we introduce a physics-driven neural retargeting algorithm that translates large-scale motion capture to humanoid embodiments while preserving physical plausibility for contact-rich interactions. Second, we learn a unified multimodal controller that supports both dense references and sparse task specifications, under sensing ranging from accurate motion-capture state to noisy egocentric visual inputs. We distill a universal tracking policy into this controller, compress motor skills into a compact latent space, and apply reinforcement learning finetuning to expand coverage and improve robustness under out-of-distribution scenarios. This enables coordinated whole-body behavior from sparse intent without test-time reference motions. We evaluate ULTRA in simulation and on a real Unitree G1 humanoid. Results show that ULTRA generalizes to autonomous, goal-conditioned whole-body loco-manipulation from egocentric perception, consistently outperforming tracking-only baselines with limited skills.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。