用多尺度与注意力优化,让CLIP模型零样本分割更准
MARS-CLIP: Multi-Resolution and Attention Refined Zero-Shot Image Segmentation

- 通过多分辨率融合,同时保留细节与全局信息
- 在6个数据集上平均性能超越现有方法
- 适合需要零样本分割的视觉任务研究者
对比语言-图像预训练(CLIP)在零样本迁移中表现出色,但在密集预测任务中常因空间分辨率低和结构信息丢失而表现不佳。为此,我们提出MARS-CLIP(多分辨率与注意力精炼的CLIP分割框架),一种新型零样本语义分割方法。该方法引入两个关键策略:(i) 多分辨率特征提取模块,融合局部细粒度特征与全局上下文,以突破输入分辨率限制;(ii) 注意力精炼机制,将中间层的空间与颜色先验注入最终自注意力块,精准恢复物体边界。在六个公开数据集上的实验表明,MARS-CLIP显著优于现有最先进方法。
原文摘要 · Abstract (English)
Contrastive Language-Image Pre-training (CLIP) has demonstrated impressive capabilities in zero-shot transfer but often struggles with dense prediction tasks due to low spatial resolution and the loss of structural information. To address these limitations, we propose MARS-CLIP (Multi-resolution and Attention Refined Segmentation for CLIP), a novel framework for zero-shot semantic segmentation. Our approach introduces two key strategies: (i) a multi-resolution feature extraction module that fuses local fine-grained features with global context to overcome input resolution constraints, and (ii) an attention refinement mechanism that injects spatial and color biases from intermediate layers into the final self-attention block to accurately restore object boundaries. A set of experiments on six public datasets demonstrates that MARS-CLIP significantly outperforms state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。