arXiv:2511.07238cs.CVcs.AI2025-11被引 2

用语言提示增强视觉模型,让自动驾驶更可靠识别未知障碍物。

Leveraging Text-Driven Semantic Variation for Robust OOD Segmentation

  • 用文本提示引导模型学习多样化的语义特征
  • 在多个数据集上实现像素级与物体级的最优表现
  • 适合提升自动驾驶系统对未知路况的鲁棒性

在自动驾驶与机器人领域,保障道路安全和可靠决策关键在于对分布外(OOD)物体的分割。尽管已有大量方法用于检测道路上的异常物体,但利用蕴含丰富语言知识的视觉-语言空间仍属未充分探索领域。本文假设,在真实自动驾驶复杂场景中引入语言线索尤为有益。为此,提出一种新方法:训练一个文本驱动的OOD分割模型,使其在视觉-语言空间中学习语义多样的对象表征。具体而言,结合视觉-语言编码器与Transformer解码器,采用位于不同语义距离处的基于距离的OOD提示,并引入OOD语义增强来优化分布外表示。通过对齐视觉与文本信息,该方法能有效泛化至未见物体,在多样驾驶环境中实现稳健的OOD分割。我们在Fishyscapes、Segment-Me-If-You-Can和Road Anomaly等公开数据集上进行大量实验,结果表明该方法在像素级与物体级评估中均达到当前最优性能,凸显了基于视觉-语言的OOD分割在提升未来自动驾驶系统安全性与可靠性方面的潜力。

原文摘要 · Abstract (English)

In autonomous driving and robotics, ensuring road safety and reliable decision-making critically depends on out-of-distribution (OOD) segmentation. While numerous methods have been proposed to detect anomalous objects on the road, leveraging the vision-language space-which provides rich linguistic knowledge-remains an underexplored field. We hypothesize that incorporating these linguistic cues can be especially beneficial in the complex contexts found in real-world autonomous driving scenarios. To this end, we present a novel approach that trains a Text-Driven OOD Segmentation model to learn a semantically diverse set of objects in the vision-language space. Concretely, our approach combines a vision-language model's encoder with a transformer decoder, employs Distance-Based OOD prompts located at varying semantic distances from in-distribution (ID) classes, and utilizes OOD Semantic Augmentation for OOD representations. By aligning visual and textual information, our approach effectively generalizes to unseen objects and provides robust OOD segmentation in diverse driving environments. We conduct extensive experiments on publicly available OOD segmentation datasets such as Fishyscapes, Segment-Me-If-You-Can, and Road Anomaly datasets, demonstrating that our approach achieves state-of-the-art performance across both pixel-level and object-level evaluations. This result underscores the potential of vision-language-based OOD segmentation to bolster the safety and reliability of future autonomous driving systems.

OOD分割视觉语言自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。