arXiv:2609.09158cs.ROcs.AI2026-09

TANGO让机器人像人一样在复杂环境里边看边走,靠语言指令指挥全身动作。

TANGO: Humanoid Navigation in Cluttered Environments with a Whole-Body Vision-Language-Action Model

论文配图:TANGO: Humanoid Navigation in Cluttered Environments with a Whole-Body Vision-Language-Action Model
图 1 · 摘自论文原文
  • 用视觉-语言-动作一体化模型,直接输出29自由度全身动作
  • 仿真训练生成多样无碰撞行为,实现语言引导的全身运动控制
  • 零样本部署到真实机器人,在未训练场景中稳健导航

我们研究了人形机器人在杂乱室内环境中的导航问题。与传统将导航建模为二维路径规划的方法不同,人形机器人在复杂3D空间中穿行需持续进行几何感知的全身协调,包括手臂位置、躯干姿态和步态的动态调整以避免碰撞。为此,我们提出TANGO,首个面向语言条件下的复杂环境人形机器人全身导航框架。给定自然语言指令和第一人称RGB观测,TANGO直接预测29自由度关节空间动作,用于下游全身控制。TANGO完全在仿真中训练,通过全局路径规划、运动学全身运动生成、障碍物感知的动作编辑以及基于强化学习的跟踪,合成多样化无碰撞的通行行为。该流程为学习语言驱动的全身策略提供了动态可行的动作监督。大量仿真实验表明,TANGO在视觉语言导航任务中达到最先进水平,且在需障碍物协商的复杂场景中显著优于强基线模块化方法。最后,我们将TANGO零样本部署于Unitree G1人形机器人,在真实杂乱环境中实现了稳健的语言引导通行,全程未使用任何真实导航数据。

原文摘要 · Abstract (English)

We study the problem of navigating cluttered indoor environments with a humanoid robot. Unlike conventional methods that model navigation as a 2D path planning problem, humanoid traversal in cluttered environments requires continuous geometry-aware whole-body adaptation, including coordinated arm placement, torso adjustment, and gait modulation for collision-free movement through complex 3D spaces. We introduce TANGO, the first whole-body vision-language navigation framework for language-conditioned humanoid traversal in cluttered environments. Given a natural-language instruction and egocentric RGB observations, TANGO directly predicts 29-DoF joint-space actions for downstream whole-body control. We train TANGO entirely in simulation by synthesizing diverse collision-free traversal behaviors via global path planning, kinematic whole-body motion generation, obstacle-aware motion editing, and RL-based tracking. This pipeline provides dynamically feasible action supervision for learning language-conditioned whole-body policies. In extensive simulation experiments, TANGO demonstrates state-of-the-art performance in vision-language navigation, while outperforming strong modular baselines in navigating challenging scenes requiring obstacle negotiation. Lastly, we deploy TANGO zero-shot on a Unitree G1 humanoid robot, and observe robust language-guided traversal in cluttered real-world scenes without training on any real-world navigation data.

人形机器人视觉语言导航全身控制零样本部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。