arXiv:2504.00356cs.CVcs.AI2025-04CVPR被引 16

不训练模型,用全局局部融合提升零样本图像指代分割精度

Hybrid Global-Local Representation with Augmented Spatial Guidance for Zero-Shot Referring Image Segmentation

  • 结合掩码特有细节与周围上下文信息,增强区域表征
  • 引入空间引导增强策略,显著提升定位准确性
  • 无需训练,适合快速部署到跨模态应用中

近期基于SAM和CLIP等模型的零样本指代图像分割(RIS)取得显著进展,实现了视觉与文本信息的有效对齐。然而,精确且高质量的掩码区域表征仍是关键挑战,制约了RIS性能的进一步提升。本文提出一种无需训练的混合全局-局部特征提取方法,将掩码特定细节与周围上下文信息相结合,强化区域表征能力。为进一步增强掩码区域与指代表达间的对齐,我们设计了一种空间引导增强策略,通过整合多源空间线索,提升空间一致性,从而实现更鲁棒、精准的指代分割。在标准RIS基准上的大量实验表明,该方法显著优于现有零样本RIS模型,性能实现显著提升。代码已开源:https://github.com/fhgyuanshen/HybridGL。

原文摘要 · Abstract (English)

Recent advances in zero-shot referring image segmentation (RIS), driven by models such as the Segment Anything Model (SAM) and CLIP, have made substantial progress in aligning visual and textual information. Despite these successes, the extraction of precise and high-quality mask region representations remains a critical challenge, limiting the full potential of RIS tasks. In this paper, we introduce a training-free, hybrid global-local feature extraction approach that integrates detailed mask-specific features with contextual information from the surrounding area, enhancing mask region representation. To further strengthen alignment between mask regions and referring expressions, we propose a spatial guidance augmentation strategy that improves spatial coherence, which is essential for accurately localizing described areas. By incorporating multiple spatial cues, this approach facilitates more robust and precise referring segmentation. Extensive experiments on standard RIS benchmarks demonstrate that our method significantly outperforms existing zero-shot RIS models, achieving substantial performance gains. We believe our approach advances RIS tasks and establishes a versatile framework for region-text alignment, offering broader implications for cross-modal understanding and interaction. Code is available at https://github.com/fhgyuanshen/HybridGL .

零样本分割跨模态对齐图像理解无训练方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。