提出新模型让视觉系统在复杂场景中准确识别物体类别。
Open-RGBT: Open-vocabulary RGB-T Zero-shot Semantic Segmentation in Open-world Environments
- 用视觉提示增强类别理解,生成更准的检测框。
- 结合CLIP模型提升图文匹配度,减少分类歧义。
- 适合开放世界下实时感知任务,尤其对热成像场景有效。
语义分割是实现有效场景理解的关键技术。传统RGB-T语义分割模型因依赖预训练模型和固定类别,在多样场景中泛化能力差。近年来,视觉语言模型(VLMs)推动了从封闭集向开放词汇语义分割的转变。然而,这些模型在复杂场景中表现受限,主要源于可见光与热成像模态间的异质性。为此,我们提出Open-RGBT,一种新型开放词汇RGB-T语义分割模型。通过引入视觉提示获取实例级检测建议,增强类别理解;同时利用CLIP模型评估图像-文本相似性,提升语义一致性,缓解类别识别模糊问题。实证结果表明,Open-RGBT在多样化且挑战性的现实场景中表现优异,甚至在野外环境中也显著优于现有方法,大幅推进了RGB-T语义分割的发展。
原文摘要 · Abstract (English)
Semantic segmentation is a critical technique for effective scene understanding. Traditional RGB-T semantic segmentation models often struggle to generalize across diverse scenarios due to their reliance on pretrained models and predefined categories. Recent advancements in Visual Language Models (VLMs) have facilitated a shift from closed-set to open-vocabulary semantic segmentation methods. However, these models face challenges in dealing with intricate scenes, primarily due to the heterogeneity between RGB and thermal modalities. To address this gap, we present Open-RGBT, a novel open-vocabulary RGB-T semantic segmentation model. Specifically, we obtain instance-level detection proposals by incorporating visual prompts to enhance category understanding. Additionally, we employ the CLIP model to assess image-text similarity, which helps correct semantic consistency and mitigates ambiguities in category identification. Empirical evaluations demonstrate that Open-RGBT achieves superior performance in diverse and challenging real-world scenarios, even in the wild, significantly advancing the field of RGB-T semantic segmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。