arXiv:2605.11939cs.CV2026-05

提升视觉语言模型在长尾数据上的分类能力,尤其增强弱类别区分度。

Cluster-Aware Neural Collapse Prompt Tuning for Long-Tailed Generalization of Vision-Language Models

论文配图:Cluster-Aware Neural Collapse Prompt Tuning for Long-Tailed Generalization of Vision-Language Models
图 1 · 摘自论文原文
  • 构建簇不变空间,约束提示调优的局部结构以保护全局语义。
  • 引入三重损失优化,实现类内紧凑、类间分离,提升尾部类别辨识度。
  • 适用于长尾分布数据,特别适合需兼顾泛化与细粒度分类的场景。

提示学习已成为替代微调预训练视觉语言模型(VLMs)的有效方法。尽管前景广阔,现有方法在适应类别不平衡数据集时仍难以维持尾部类别的判别能力。本文提出簇感知神经坍缩提示调优(CPT),在不牺牲整体泛化性能的前提下,增强提示调优后VLM对尾部类别的判别性。首先,通过挖掘预训练VLM的语义分配并映射到提示调优特征,构建簇不变空间,计算簇级边界,并将约束限制在局部邻域,减少对预训练VLM全局语义结构的干扰。其次,引入由神经坍缩驱动的判别性优化,包含文本等角紧框架(ETF)分离损失、类内收敛损失和旋转稳定性损失,协同作用以优化类内几何结构,实现更好的类间分离与类内对齐。在11个多样化数据集上的大量实验表明,CPT优于现有最先进方法,在长尾类别上表现更强,且对未见类别具有良好的泛化能力。

原文摘要 · Abstract (English)

Prompt learning has emerged as an efficient alternative to fine-tuning pre-trained vision-language models (VLMs). Despite its promise, current methods still struggle to maintain tail-class discriminability when adapting to class-imbalanced datasets. In this work, we propose cluster-aware neural collapse prompt tuning (CPT), which enhances the discriminability of tail classes in prompt-tuned VLMs without sacrificing their overall generalization. First, we design a cluster-invariant space by mining semantic assignments from the pre-trained VLM and mapping them to prompt-tuned features. This computes cluster-level boundaries and restricts the constraints to local neighborhoods, which reduces interference with the global semantic structure of the pre-trained VLM. Second, we introduce neural-collapse-driven discriminability optimization with three losses: textual Equiangular Tight Frame (ETF) separation loss, class-wise convergence loss, and rotation stabilization loss. These losses work together to shape intra-cluster geometry for better inter-class separation and intra-class alignment. Extensive experiments on 11 diverse datasets demonstrate that CPT outperforms SOTA methods, with stronger performance on long-tail classes and good generalization to unseen classes.

提示调优长尾分布视觉语言模型类间分离

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。