提出动态3D视觉语言规划模型,实现可解释的智能体导航与任务执行。
D3D-VLP: Dynamic 3D Vision-Language-Planning Model for Embodied Grounding and Navigation
- 构建统一3D思维链框架,融合规划、定位、导航与问答
- 在1000万样本数据上达成多个任务最优表现
- 适合需要真实世界交互的机器人系统研发
具身智能体面临端到端模型缺乏可解释性与显式三维推理,而模块化系统忽略组件间关联性的困境。为此,我们提出动态3D视觉-语言-规划模型(D3D-VLP)。该模型引入两项关键创新:1)动态3D思维链(3D CoT),将规划、定位、导航与问答统一于单一3D-VLM与思维链流程;2)碎片化监督协同学习(SLFS)策略,采用掩码自回归损失,从海量且部分标注的混合数据中学习,使各思维链组件相互增强并隐式监督。我们构建了包含1000万条混合样本的大规模数据集,源自5000个真实扫描与20000个合成场景,支持在线学习方法如强化学习与DAgger。D3D-VLP在多个基准上取得最先进结果,包括视觉语言导航(R2R-CE、REVERIE-CE、NavRAG-CE)、目标导向导航(HM3D-OVON)及任务导向序列定位与导航(SG3D)。真实世界移动操作实验进一步验证其有效性。
原文摘要 · Abstract (English)
Embodied agents face a critical dilemma that end-to-end models lack interpretability and explicit 3D reasoning, while modular systems ignore cross-component interdependencies and synergies. To bridge this gap, we propose the Dynamic 3D Vision-Language-Planning Model (D3D-VLP). Our model introduces two key innovations: 1) A Dynamic 3D Chain-of-Thought (3D CoT) that unifies planning, grounding, navigation, and question answering within a single 3D-VLM and CoT pipeline; 2) A Synergistic Learning from Fragmented Supervision (SLFS) strategy, which uses a masked autoregressive loss to learn from massive and partially-annotated hybrid data. This allows different CoT components to mutually reinforce and implicitly supervise each other. To this end, we construct a large-scale dataset with 10M hybrid samples from 5K real scans and 20K synthetic scenes that are compatible with online learning methods such as RL and DAgger. Our D3D-VLP achieves state-of-the-art results on multiple benchmarks, including Vision-and-Language Navigation (R2R-CE, REVERIE-CE, NavRAG-CE), Object-goal Navigation (HM3D-OVON), and Task-oriented Sequential Grounding and Navigation (SG3D). Real-world mobile manipulation experiments further validate the effectiveness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。