arXiv:2511.15967cs.CV2025-11AAAI被引 3

用信息论方法稳定CLIP微调时的图文对齐,提升开放词汇语义分割性能

InfoCLIP: Bridging Vision-Language Pretraining and Open-Vocabulary Semantic Segmentation via Information-Theoretic Alignment Transfer

  • 通过信息论约束压缩图文对齐特征,减少噪声
  • 最大化预训练与微调模型间的互信息,传递紧凑语义关系
  • 适合需要稳定微调且关注开放词汇分割的研究者

近期,CLIP强大的泛化能力推动了开放词汇语义分割的发展,可使用任意文本标签像素。然而,现有方法在有限类别上微调CLIP时易过拟合,并破坏预训练的视觉-语言对齐。为稳定微调过程中的模态对齐,我们提出InfoCLIP,从信息论角度将预训练CLIP的对齐知识迁移至分割任务。具体而言,该迁移由两个基于互信息的新目标引导:首先,压缩预训练CLIP中像素-文本对齐特征,以降低其在图像-文本监督下学习到的粗粒度局部语义表示带来的噪声;其次,最大化预训练模型与微调模型之间对齐知识的互信息,以传递适配分割任务的紧凑局部语义关系。在多个基准上的广泛评估验证了InfoCLIP在增强CLIP微调于开放词汇语义分割中的有效性,展示了其在非对称迁移中的适应性与优越性。

原文摘要 · Abstract (English)

Recently, the strong generalization ability of CLIP has facilitated open-vocabulary semantic segmentation, which labels pixels using arbitrary text. However, existing methods that fine-tune CLIP for segmentation on limited seen categories often lead to overfitting and degrade the pretrained vision-language alignment. To stabilize modality alignment during fine-tuning, we propose InfoCLIP, which leverages an information-theoretic perspective to transfer alignment knowledge from pretrained CLIP to the segmentation task. Specifically, this transfer is guided by two novel objectives grounded in mutual information. First, we compress the pixel-text modality alignment from pretrained CLIP to reduce noise arising from its coarse-grained local semantic representations learned under image-text supervision. Second, we maximize the mutual information between the alignment knowledge of pretrained CLIP and the fine-tuned model to transfer compact local semantic relations suited for the segmentation task. Extensive evaluations across various benchmarks validate the effectiveness of InfoCLIP in enhancing CLIP fine-tuning for open-vocabulary semantic segmentation, demonstrating its adaptability and superiority in asymmetric transfer.

语义分割CLIP信息论开放词汇

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。