构建多场景语言导航数据集,让机器人按指令找物并停在1米内。
MiniVLA-Nav v1: A Multi-Scene Simulation Dataset for Language-Conditioned Robot Navigation

- 基于真实感仿真环境生成带视觉与动作标签的导航数据
- 1174个任务中机器人需在3.5-7米距离内准确抵达目标物体
- 支持跨模板、跨类别泛化测试,适合训练语言导航模型
我们提出MiniVLA-Nav v1,一个用于语言条件物体接近(LCOA)导航的仿真数据集:给定简短自然语言指令,NVIDIA Nova Carter差速驱动机器人需在四个逼真Isaac Sim环境(办公室、医院、全仓库、多货架仓库)中导航至指定物体,并在1米范围内停止。每个包含1,174个场景的片段配以同步的640x640 RGB图像、度量深度图(float32,单位米)和实例分割掩码,同时记录每秒60帧的连续动作(v, omega)及7x7分词专家动作标签。通过三种起始距离层级(近:1.5-3.5米,中:3.5-7.0米,远:全局精选点;起始距离与轨迹长度相关性Pearson r=0.94)、12类物体、18个训练模板和12个改写式分布外模板确保轨迹多样性。五个评估划分支持分布内准确率、模板改写鲁棒性与分布外物体类别基准测试。数据集已公开于https://huggingface.co/datasets/alibustami/miniVLA-Nav
原文摘要 · Abstract (English)
We present MiniVLA-Nav v1, a simulation dataset for Language-Conditioned Object Approach (LCOA) navigation: given a short natural-language instruction, an NVIDIA Nova Carter differential-drive robot must navigate to the named object and stop within 1 m across four photorealistic Isaac Sim environments (Office, Hospital, Full Warehouse, and Warehouse with Multiple Shelves). Each of the 1,174 episodes pairs an instruction with synchronized 640x640 RGB images, metric depth maps (float32, metres), and instance segmentation masks, together with continuous (v,omega) and 7x7 tokenized expert action labels recorded at 60 Hz from a vision-based proportional controller. Trajectory diversity is ensured through three spawn-distance tiers (near: 1.5-3.5 m, mid: 3.5-7.0 m, far: global curated points; Pearson r=0.94 between spawn distance and trajectory length), 12 object categories, 18 training templates, and 12 paraphrase-OOD templates. Five evaluation splits support in-distribution accuracy, template-paraphrase robustness, and OOD object-category benchmarking. The dataset is publicly available at https://huggingface.co/datasets/alibustami/miniVLA-Nav
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。