arXiv:2411.13578cs.CVcs.AI2024-11被引 8

无需重新训练,就能精准识别多标签场景下的异常样本。

CMOOD: Concept-based Multi-label OOD Detection

  • 基于概念扩展增强视觉语言模型的语义空间。
  • 在VOC和COCO上平均AUROC达95%。
  • 适合处理标签间有依赖关系的真实多标签任务。

现有方法在复杂多标签场景下难以捕捉语义关系与标签共现模式,通常需大量训练数据且无法泛化到未见标签组合。尽管大语言模型推动了零样本异常检测发展,但主要针对单标签场景。为此,本文提出COOD框架,利用预训练视觉-语言模型,结合概念驱动的标签扩展策略与新型评分函数。通过为每个标签引入正负概念丰富语义空间,模型能有效建模复杂标签依赖,无需额外训练即可精确区分异常样本。大量实验表明,该方法在VOC和COCO数据集上平均AUROC达约95%,在不同标签数量和异常类型下均保持鲁棒性能。

原文摘要 · Abstract (English)

How can models effectively detect out-of-distribution (OOD) samples in complex, multi-label settings without extensive retraining? Existing OOD detection methods struggle to capture the intricate semantic relationships and label co-occurrences inherent in multi-label settings, often requiring large amounts of training data and failing to generalize to unseen label combinations. While large language models have revolutionized zero-shot OOD detection, they primarily focus on single-label scenarios, leaving a critical gap in handling real-world tasks where samples can be associated with multiple interdependent labels. To address these challenges, we introduce COOD, a novel zero-shot multi-label OOD detection framework. COOD leverages pre-trained vision-language models, enhancing them with a concept-based label expansion strategy and a new scoring function. By enriching the semantic space with both positive and negative concepts for each label, our approach models complex label dependencies, precisely differentiating OOD samples without the need for additional training. Extensive experiments demonstrate that our method significantly outperforms existing approaches, achieving approximately 95% average AUROC on both VOC and COCO datasets, while maintaining robust performance across varying numbers of labels and different types of OOD samples.

多标签异常检测零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。