arXiv:2412.16609cs.CV2024-12被引 1

用语义概念提升多图共显著目标检测精度

Concept Guided Co-salient Object Detection

  • 引入共享文本概念提供高层语义引导
  • 在三个基准数据集上达到新高,鲁棒性更强
  • 适合需要跨图像语义理解的视觉任务

共显著目标检测(Co-SOD)旨在识别一组相关图像中的共同显著对象。现有方法多依赖低层视觉模式,缺乏语义先验,限制了检测性能。本文提出ConceptCoSOD框架,通过从输入图像组中提取共享的文本概念,引入高层语义知识以指导检测过程。为提升概念质量,分析扩散步数影响并设计重采样策略,选取更具信息量的步骤学习鲁棒概念。该语义先验与增强表示相结合,使模型在复杂视觉条件下仍能实现精准一致的分割。在三个基准数据集及五种损坏设置下的大量实验表明,ConceptCoSOD显著优于现有方法,在准确率和泛化能力上均有提升。

原文摘要 · Abstract (English)

Co-salient object detection (Co-SOD) aims to identify common salient objects across a group of related images. While recent methods have made notable progress, they typically rely on low-level visual patterns and lack semantic priors, limiting their detection performance. We propose ConceptCoSOD, a concept-guided framework that introduces high-level semantic knowledge to enhance co-saliency detection. By extracting shared text-based concepts from the input image group, ConceptCoSOD provides semantic guidance that anchors the detection process. To further improve concept quality, we analyze the effect of diffusion timesteps and design a resampling strategy that selects more informative steps for learning robust concepts. This semantic prior, combined with the resampling-enhanced representation, enables accurate and consistent segmentation even in challenging visual conditions. Extensive experiments on three benchmark datasets and five corrupted settings demonstrate that ConceptCoSOD significantly outperforms existing methods in both accuracy and generalization.

共显著检测语义引导扩散模型图像分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。