arXiv:2606.19088cs.RO2026-06

让机器人理解语言指令时更准确定位物体位置。

ReSiReg: Towards Spatially Consistent Semantics in Language-Conditioned Robotic Tasks

论文配图:ReSiReg: Towards Spatially Consistent Semantics in Language-Conditioned Robotic Tasks
图 1 · 摘自论文原文
  • 用视觉原型重建特征,提升语言-视觉嵌入的空间一致性。
  • 在OVSS和3D映射任务中,密集检索准确率显著提升。
  • 模型仅25M参数,适合部署在资源受限的机器人上。

视觉语言模型(VLMs)使机器人能够遵循开放语言指令,但其密集嵌入存在噪声且缺乏空间一致性,这对需要同时进行语义与三维空间推理的机器人应用构成挑战。本文分析了近期VLMs的空间结构,提出ReSiReg:一种利用空间一致的中间特征进行重建的方法。该方法将中间特征聚类为视觉原型,提取其语言描述,并将每个图像块重构为原型级语言嵌入的软混合。在多种骨干网络上,于OVSS和3D映射任务中进行了定量评估,结果表明密集检索性能显著提升;真实操作场景中目标激活区域更具空间一致性。此外,本文还提供了一个仅25M参数的紧凑型密集VLM,体积远小于且性能媲美ViT-B基线模型。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) enable robots to follow open-language instructions. However, dense VLM embeddings have shown to be noisy and lack spatial consistency. This is problematic for robotic applications, which require simultaneous reasoning over semantics and 3D space. We examine spatial structure across recent VLMs and propose ReSiReg, a feature reconstruction method that uses spatially consistent VLM intermediates to improve dense language-grounded retrieval. ReSiReg clusters intermediates into visual prototypes, derives their language descriptors, and reconstructs each patch as a soft mixture of prototype-level language embeddings. We evaluate quantitatively on OVSS and 3D mapping across backbones, and qualitatively in real-world manipulation scenes. Quantitative results show improved dense retrieval; manipulation scenes show more spatially consistent target activations. We further provide a compact 25M dense VLM for robotic applications, substantially smaller than and competitive with ViT-B baselines. Available at https://resireg.github.io

机器人视觉语言空间一致性轻量化模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。