arXiv:2503.24065cs.CVcs.RO2025-03ICCV被引 15

用轻量模块提升视觉语言导航性能,降低计算开销。

COSMO: Combination of Selective Memorization for Low-cost Vision-and-Language Navigation

  • 融合状态空间与Transformer,设计双定制模块增强跨模态交互。
  • 在三个主流数据集上表现优异,计算成本显著降低。
  • 适合资源受限场景下的智能导航系统部署。

视觉-语言导航(VLN)任务因在家庭助理等领域的应用潜力而受到关注。现有方法虽基于Transformer架构并引入外部知识库或地图信息以提升性能,但导致模型增大、计算成本升高。本文提出新型架构COSMO(COmbination of Selective MemOrization),融合状态空间模块与Transformer模块,并引入两种专用于VLN的可选择状态空间模块:环形选择扫描(RSS)与跨模态选择状态空间模块(CS3)。RSS实现单次扫描内的完整跨模态交互,CS3将选择性状态空间模块适配为双流结构,增强跨模态信息获取。在三个主流VLN基准测试REVERIE、R2R和R2R-CE上的实验表明,所提模型不仅达到竞争性导航性能,还显著降低计算成本。

原文摘要 · Abstract (English)

Vision-and-Language Navigation (VLN) tasks have gained prominence within artificial intelligence research due to their potential application in fields like home assistants. Many contemporary VLN approaches, while based on transformer architectures, have increasingly incorporated additional components such as external knowledge bases or map information to enhance performance. These additions, while boosting performance, also lead to larger models and increased computational costs. In this paper, to achieve both high performance and low computational costs, we propose a novel architecture with the COmbination of Selective MemOrization (COSMO). Specifically, COSMO integrates state-space modules and transformer modules, and incorporates two VLN-customized selective state space modules: the Round Selective Scan (RSS) and the Cross-modal Selective State Space Module (CS3). RSS facilitates comprehensive inter-modal interactions within a single scan, while the CS3 module adapts the selective state space module into a dual-stream architecture, thereby enhancing the acquisition of cross-modal interactions. Experimental validations on three mainstream VLN benchmarks, REVERIE, R2R, and R2R-CE, not only demonstrate competitive navigation performance of our model but also show a significant reduction in computational costs.

视觉语言导航轻量模型状态空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。