arXiv:2604.15946cs.CVcs.RO2026-04

用双目视觉提升开放词汇语义分割精度,解决遮挡与边界模糊问题。

SENSE: Stereo OpEN Vocabulary SEmantic Segmentation

论文配图:SENSE: Stereo OpEN Vocabulary SEmantic Segmentation
图 1 · 摘自论文原文
  • 融合双目图像与视觉语言模型,利用几何信息增强空间推理。
  • 在PhraseStereo上比基线提升2.9%平均精度,零样本下表现优异。
  • 适合自动驾驶与智能交通系统,支持自然语言驱动的场景理解。

开放词汇语义分割使模型能识别固定类别之外的物体或图像区域,适用于动态环境。然而,现有方法多依赖单视角图像,在遮挡和近物边界处常出现空间精度不足的问题。本文提出SENSE,首个基于双目视觉的开放词汇语义分割方法,结合立体视觉与视觉-语言模型,引入几何线索以提升空间推理与分割准确性。在PhraseStereo数据集上训练后,该方法在短语定位任务中表现强劲,并展现出出色的零样本泛化能力。在PhraseStereo上,相比基线方法提升2.9%平均精度,较最优竞争方法提升0.76%;在Cityscapes上相对基线提升3.5% mIoU,KITTI上提升18%。通过联合语义与几何推理,SENSE可实现由自然语言驱动的精准场景理解,对自主机器人与智能交通系统具有重要意义。

原文摘要 · Abstract (English)

Open-vocabulary semantic segmentation enables models to segment objects or image regions beyond fixed class sets, offering flexibility in dynamic environments. However, existing methods often rely on single-view images and struggle with spatial precision, especially under occlusions and near object boundaries. We propose SENSE, the first work on Stereo OpEN Vocabulary SEmantic Segmentation, which leverages stereo vision and vision-language models to enhance open-vocabulary semantic segmentation. By incorporating stereo image pairs, we introduce geometric cues that improve spatial reasoning and segmentation accuracy. Trained on the PhraseStereo dataset, our approach achieves strong performance in phrase-grounded tasks and demonstrates generalization in zero-shot settings. On PhraseStereo, we show a +2.9% improvement in Average Precision over the baseline method and +0.76% over the best competing method. SENSE also provides a relative improvement of +3.5% mIoU on Cityscapes and +18% on KITTI compared to the baseline work. By jointly reasoning over semantics and geometry, SENSE supports accurate scene understanding from natural language, essential for autonomous robots and Intelligent Transportation Systems.

语义分割双目视觉开放词汇自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。