arXiv:2603.27577cs.CVcs.RO2026-03被引 4

用结构化语言描述视觉信息,让导航模型更小更强、适应新环境。

Structured Observation Language for Efficient and Generalizable Vision-Language Navigation

  • 将视角图像分块提取语义、颜色、深度信息,转为结构化文本
  • 在R2R和RxR上模型更小、训练数据少,跨环境泛化能力显著提升
  • 适合追求轻量化与强泛化的视觉语言导航研究者

视觉-语言导航(VLN)要求智能体根据自然语言指令在复杂环境中导航,通常需要紧密融合视觉与语言模态。现有方法常将原始图像转化为视觉标记或隐式特征,依赖大规模视觉预训练,在光照、纹理等环境变化下泛化能力差。为此,我们提出SOL-Nav(Structured Observation Language for Navigation),一种将第一人称视觉观测转换为紧凑结构化语言描述的新框架,实现高效且可泛化的导航。具体而言,将RGB-D图像划分为NxN网格,对每个网格单元提取语义、颜色和深度代表性信息,生成结构化文本,并与语言指令拼接后作为纯语言输入给预训练语言模型(PLM)。在标准VLN基准(R2R、RxR)及真实世界部署的实验表明,SOL-Nav显著降低模型规模与训练数据依赖,充分释放PLM的推理与表征能力,在未见环境中表现优异。

原文摘要 · Abstract (English)

Vision-Language Navigation (VLN) requires an embodied agent to navigate complex environments by following natural language instructions, which typically demands tight fusion of visual and language modalities. Existing VLN methods often convert raw images into visual tokens or implicit features, requiring large-scale visual pre-training and suffering from poor generalization under environmental variations (e.g., lighting, texture). To address these issues, we propose SOL-Nav (Structured Observation Language for Navigation), a novel framework that translates egocentric visual observations into compact structured language descriptions for efficient and generalizable navigation. Specifically, we divide RGB-D images into a NxN grid, extract representative semantic, color, and depth information for each grid cell to form structured text, and concatenate this with the language instruction as pure language input to a pre-trained language model (PLM). Experimental results on standard VLN benchmarks (R2R, RxR) and real-world deployments demonstrate that SOL-Nav significantly reduces the model size and training data dependency, fully leverages the reasoning and representation capabilities of PLMs, and achieves strong generalization to unseen environments.

视觉导航多模态语言模型轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。