arXiv:2512.10262cs.CV2025-12

用视觉语言模型提升未知类别发现准确率,尤其擅长处理数据不均衡问题。

VLM-NCD:Novel Class Discovery with Vision-Based Large Language Models

  • 融合图像文本语义,用原型引导聚类发现新类别。
  • 在CIFAR-100上对未知类别的识别准确率提升25.3%。
  • 首次在NCD任务中展现对长尾分布的强鲁棒性,适合实际数据场景。

新型类别发现旨在利用已知类别的先验知识,从无标签数据中分类并发现未知类别。现有图像类NCD方法主要依赖视觉特征,存在特征区分度不足和数据长尾分布等问题。本文提出LLM-NCD,一种基于多模态融合的框架,通过结合视觉-文本语义与原型引导聚类,突破上述瓶颈。核心创新在于联合优化已知类别的图像与文本特征,构建其簇中心与语义原型,并采用双阶段发现机制,通过语义亲和阈值动态分离已知或未知样本,实现自适应聚类。在CIFAR-100数据集上的实验表明,相比现有方法,该方法在未知类别识别准确率上最高提升25.3%。尤为突出的是,本方法首次在NCD领域展现出对长尾分布的独特鲁棒性。

原文摘要 · Abstract (English)

Novel Class Discovery aims to utilise prior knowledge of known classes to classify and discover unknown classes from unlabelled data. Existing NCD methods for images primarily rely on visual features, which suffer from limitations such as insufficient feature discriminability and the long-tail distribution of data. We propose LLM-NCD, a multimodal framework that breaks this bottleneck by fusing visual-textual semantics and prototype guided clustering. Our key innovation lies in modelling cluster centres and semantic prototypes of known classes by jointly optimising known class image and text features, and a dualphase discovery mechanism that dynamically separates known or novel samples via semantic affinity thresholds and adaptive clustering. Experiments on the CIFAR-100 dataset show that compared to the current methods, this method achieves up to 25.3% improvement in accuracy for unknown classes. Notably, our method shows unique resilience to long tail distributions, a first in NCD literature.

类别发现多模态长尾分布大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。