arXiv:2505.03361cs.CV2025-05被引 1

用无限概念提升零样本学习可解释性,解决大模型幻觉问题

Interpretable Zero-shot Learning with Infinite Class Concepts

  • 用大模型动态生成海量短语级类别概念
  • 通过熵评分筛选出最可迁移、可区分的概念
  • 在3个数据集上性能显著提升,结果直观可解释

零样本学习(ZSL)旨在通过图像与中间语义的对齐来识别未见类别,传统方法依赖人工标注概念或类别定义。近期研究利用大语言模型(LLM)自动生成类别文档,但常面临分类过程不透明及大模型幻觉问题,导致生成非视觉语义。本文重新定义ZSL中的类别语义,强调可迁移性与可区分性,提出零样本学习无限类别概念框架(InfZSL)。该方法利用大模型动态生成无限数量的短语级类别概念,并引入基于熵的评分机制与“优劣”概念选择策略,确保仅保留最具可迁移性和可区分性的概念。InfZSL在三个主流基准数据集上表现显著优于现有方法,且生成的概念高度可解释、与图像强关联。代码将在论文接受后发布。

原文摘要 · Abstract (English)

Zero-shot learning (ZSL) aims to recognize unseen classes by aligning images with intermediate class semantics, like human-annotated concepts or class definitions. An emerging alternative leverages Large-scale Language Models (LLMs) to automatically generate class documents. However, these methods often face challenges with transparency in the classification process and may suffer from the notorious hallucination problem in LLMs, resulting in non-visual class semantics. This paper redefines class semantics in ZSL with a focus on transferability and discriminability, introducing a novel framework called Zero-shot Learning with Infinite Class Concepts (InfZSL). Our approach leverages the powerful capabilities of LLMs to dynamically generate an unlimited array of phrase-level class concepts. To address the hallucination challenge, we introduce an entropy-based scoring process that incorporates a ``goodness" concept selection mechanism, ensuring that only the most transferable and discriminative concepts are selected. Our InfZSL framework not only demonstrates significant improvements on three popular benchmark datasets but also generates highly interpretable, image-grounded concepts. Code will be released upon acceptance.

零样本学习大模型可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。