构建首个面向多样化人形机器人的真实物理导航基准
HumanoidVLN: A Physics-Grounded Simulator and Benchmark for Vision-Language Navigation Across Diverse Humanoid Embodiments

- 基于物理引擎支持多款人形机器人,含10-12个下肢自由度
- 四款机器人在43.55%成功率下表现最佳,真实世界测试误差仅0.68米
- 适配主流视觉语言模型,适合研究人形机器人导航与控制协同
现有视觉语言导航(VLN)基准无法应对人形机器人的挑战:双足行走带来物理约束,不同平台的人形形态差异大,且运动导致视角畸变。我们提出HumanoidVLN,一个支持多种人形机器人配置的物理仿真平台与基准。基于NVIDIA Isaac Sim,平台包含四款机器人(Unitree G1、H1、Internal-A、Internal-B),下肢自由度10-12,身高1.17-1.80米,通过分层控制架构实现稳定行走。环境来自艺术家设计场景和3D高斯溅射重建,可通行区域超100平方米。指令由多智能体管道生成,经人工验证,共产生933条无碰撞参考轨迹,每条配有精细指令及三种风格化变体(正式、自然、随意)。在四个模型与四类机器人上,JanusVLN取得最高均值成功率达43.55%,nDTW为48.38。20次模拟到真实迁移实验中,DualVLN与Unitree G1的导航误差相关系数r=0.935,平均绝对误差0.68米,轨迹相似度达0.782(±0.188)。结果揭示了模型、控制器与人形本体间的物理交互关系。代码、数据与基准将在论文录用后公开。
原文摘要 · Abstract (English)
Vision-Language Navigation (VLN) for humanoid robots poses challenges existing benchmarks fail to address: bipedal locomotion imposes physical constraints absent from wheeled agents, humanoid morphologies vary across platforms, and egocentric observations are distorted by locomotion-induced camera dynamics. We present HumanoidVLN, a physics-grounded simulator and benchmark for VLN across diverse humanoid embodiments. Built on NVIDIA Isaac Sim, our platform supports an extensible set of humanoid configurations, demonstrated on four robots (Unitree G1, Unitree H1, Internal-A, Internal-B) spanning 10-12 lower-body DoF and heights from 1.17m to 1.80m, via a hierarchical control stack combining a reinforcement learning locomotion policy with interchangeable PD or MPC path trackers. New robots and VLN models integrate with minimal effort; we demonstrate compatibility with NaVILA, DualVLN, StreamVLN, and JanusVLN. Environments are drawn from artist-designed scenes and 3D Gaussian Splatting reconstructions, filtered for navigable areas exceeding 100 square meters. Instructions are generated by a dual generator-reviewer plus paraphraser multi-agent pipeline with human-in-the-loop verification, yielding 933 collision-aware reference episodes, each paired with one fine-grained instruction and three coarse-grained stylistic variants (formal, natural, casual). Across four models and four embodiments, JanusVLN achieves the highest mean success rate of 43.55% and nDTW of 48.38. In a 20-episode sim-to-real pilot with DualVLN and the Unitree G1, navigation errors correlate strongly (r=0.935), with a mean absolute difference of 0.68m and mean trajectory similarity of 0.782 (+/-0.188) nDTW. These results highlight the interaction between VLN models, controllers, and humanoid embodiments under physical execution. Code, benchmark, and data will be released upon acceptance at https://humanoid-vln.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。