arXiv:2506.08678cs.CV2025-06ICCV被引 3

用自蒸馏提升CLIP模型细粒度识别能力,无需额外标注。

ATAS: Any-to-Any Self-Distillation for Enhanced Open-Vocabulary Dense Prediction

论文配图:ATAS: Any-to-Any Self-Distillation for Enhanced Open-Vocabulary Dense Prediction
图 1 · 摘自论文原文
  • 通过内部自蒸馏,统一优化语义连贯性与局部对齐。
  • 在开集目标检测与分割任务上显著超越基线CLIP。
  • 仅需无标签图像,适合部署于资源受限场景。

视觉-语言模型如CLIP通过支持广泛视觉概念识别,推动了开集密集预测任务的发展。然而,CLIP在细粒度、区域级理解方面仍存在不足,限制其在密集预测任务中的表现。我们识别出两个关键因素:语义连贯性与细粒度视觉-语言对齐。现有适配方法常以牺牲语义连贯性为代价强化细粒度对齐,且依赖额外模块或监督微调。为此,我们提出任意到任意自蒸馏(ATAS),一种新方法,通过利用模型自身在所有表示层次上的知识,同时增强语义连贯性与细粒度对齐。不同于以往方法,ATAS仅使用无标签图像和内部自蒸馏过程,优化CLIP视觉编码器的表示,保持局部语义一致性的同时提升局部细节识别能力。在开集目标检测与语义分割基准测试中,ATAS取得显著性能提升,优于基线CLIP模型。结果验证了该方法的有效性,并强调了联合保持语义连贯性与细粒度对齐对先进开集密集预测的重要性。

原文摘要 · Abstract (English)

Vision-language models such as CLIP have recently propelled open-vocabulary dense prediction tasks by enabling recognition of a broad range of visual concepts. However, CLIP still struggles with fine-grained, region-level understanding, hindering its effectiveness on these dense prediction tasks. We identify two pivotal factors required to address this limitation: semantic coherence and fine-grained vision-language alignment. Current adaptation methods often improve fine-grained alignment at the expense of semantic coherence, and often rely on extra modules or supervised fine-tuning. To overcome these issues, we propose Any-to-Any Self-Distillation (ATAS), a novel approach that simultaneously enhances semantic coherence and fine-grained alignment by leveraging own knowledge of a model across all representation levels. Unlike prior methods, ATAS uses only unlabeled images and an internal self-distillation process to refine representations of CLIP vision encoders, preserving local semantic consistency while sharpening local detail recognition. On open-vocabulary object detection and semantic segmentation benchmarks, ATAS achieves substantial performance gains, outperforming baseline CLIP models. These results validate the effectiveness of our approach and underscore the importance of jointly maintaining semantic coherence and fine-grained alignment for advanced open-vocabulary dense prediction.

视觉语言模型自蒸馏开集学习密集预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。