将3D场景实体作为软令牌注入视觉语言模型,实现无需重训的零样本导航。
SoftNav: Injecting 3D Scene Tokens into VLMs for Embodied Navigation

- 用轻量投影器将3D物体/前沿转为软令牌直接注入VLM隐空间
- 在HM3D-OVON上达74.2%成功率,超越所有已有方法
- 零样本迁移至真实机器人、GOAT-Bench等场景,仅需1700万参数
在目标导向的具身导航中,智能体需在未见过的环境中定位指定目标,要求3D场景理解与导航推理协同工作。现有方法通过文本传递3D场景信息,但实验表明存在表征鸿沟;受控消融证实,直接嵌入层级传输显著优于文本序列化格式。本文提出SoftNav,通过轻量投影器将每个检测到的物体或前沿转化为一个连续3D表示(即软令牌),直接注入视觉语言模型(VLM)的隐藏空间。在保持3D编码器和VLM冻结的前提下,仅需约1,200个样本和约1700万可训练参数。在HM3D-OVON数据集上,SoftNav在三个划分上的成功率达74.2%/68.3%/66.7%,在成功率(SR)和路径归一化成功率(SPL)上均超越此前所有方法;相同导航策略零样本迁移至GOAT-Bench(67.2% SR)、SG3D(47.2% s-SR)及真实机器人部署,无需重新训练或修改架构。直接向VLM注入3D场景令牌可弥合表征鸿沟,实现低训练成本的可迁移导航。
原文摘要 · Abstract (English)
In goal-directed embodied navigation, where an agent must locate a specified target in an unseen environment, 3D scene understanding and navigation reasoning must work in concert. Current approaches transmit 3D scene information to vision-language models (VLMs) through text, suggesting a representation gap in our tested configurations; a controlled ablation confirms that direct embedding-level transfer significantly outperforms the evaluated text serialization formats. We introduce SoftNav, which injects entity-level 3D continuous representations -- one token per detected object or frontier -- into a VLM's hidden space as soft tokens through a lightweight projector. With the 3D encoder and VLM frozen, only ~1,200 samples and ~17M trainable parameters are needed. On HM3D-OVON, SoftNav achieves 74.2%/68.3%/66.7% SR across three splits, surpassing all prior methods in both SR and SPL; the same navigation policy transfers zero-shot to GOAT-Bench (67.2% SR), SG3D (47.2% s-SR), and real-world robot deployment without retraining or architectural modification. Injecting 3D scene tokens directly into VLMs bridges the representation gap, enabling transferable navigation with minimal training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。