arXiv:2505.16663cs.CVcs.MM2025-05被引 6

让3D文本模型帮视觉导航模型解决空间困惑,提升智能体导航精度。

CoNav: Collaborative Cross-Modal Reasoning for Embodied Navigation

  • 用3D文本模型生成空间语义知识,指导图像导航模型决策。
  • 在4个标准导航任务上显著提升成功率,路径更短(SPL更高)。
  • 适合做多模态融合、具身智能导航的研究者与开发者参考。

具身导航需要全面的场景理解与精确的空间推理。虽然图像-文本模型擅长解析像素级颜色与光照线索,3D-文本模型能捕捉体积结构与空间关系,但统一融合2D图像、3D点云与文本指令的方法面临三模态数据稀缺及模态间信念冲突的挑战。本文提出CoNav,一种协作式跨模态推理框架:预训练的3D-文本模型通过提供结构化空间-语义知识,主动引导图像-文本导航代理解决导航中的歧义。核心是跨模态信念对齐机制,仅需共享3D-文本模型生成的文本假设即可实现指导。在小规模2D-3D-文本语料上轻量微调后,导航代理学会融合视觉线索与3D-文本模型提取的空间语义信息,在四个标准具身导航基准(R2R, CVDN, REVERIE, SOON)和两个空间推理基准(ScanQA, SQA3D)上均取得显著提升。在接近导航成功率达时,路径长度(以SPL衡量)更优,展示了跨模态融合在具身导航中的潜力与挑战。

原文摘要 · Abstract (English)

Embodied navigation demands comprehensive scene understanding and precise spatial reasoning. While image-text models excel at interpreting pixel-level color and lighting cues, 3D-text models capture volumetric structure and spatial relationships. However, unified fusion approaches that jointly fuse 2D images, 3D point clouds, and textual instructions face challenges in limited availability of triple-modality data and difficulty resolving conflicting beliefs among modalities. In this work, we introduce CoNav, a collaborative cross-modal reasoning framework where a pretrained 3D-text model explicitly guides an image-text navigation agent by providing structured spatial-semantic knowledge to resolve ambiguities during navigation. Specifically, we introduce Cross-Modal Belief Alignment, which operationalizes this cross-modal guidance by simply sharing textual hypotheses from the 3D-text model to the navigation agent. Through lightweight fine-tuning on a small 2D-3D-text corpus, the navigation agent learns to integrate visual cues with spatial-semantic knowledge derived from the 3D-text model, enabling effective reasoning in embodied navigation. CoNav achieves significant improvements on four standard embodied navigation benchmarks (R2R, CVDN, REVERIE, SOON) and two spatial reasoning benchmarks (ScanQA, SQA3D). Moreover, under close navigation Success Rate, CoNav often generates shorter paths compared to other methods (as measured by SPL), showcasing the potential and challenges of fusing data from different modalities in embodied navigation. Project Page: https://oceanhao.github.io/CoNav/

具身导航跨模态融合3D理解空间推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。