arXiv:2410.22317cs.CV2024-10中稿 · WACV 2025被引 12

用多类别文本反转实现无需类别名的强分类能力

Multi-Class Textual-Inversion Secretly Yields a Semantic-Agnostic Classifier

  • 将单概念文本反转扩展为多类别,加入判别正则化提升分类性能
  • 仅需每类少量样本,12个数据集上分类与生成效果均更优
  • 适合无类别名称先验但需快速建模新类别的场景

随着大型预训练视觉语言模型(如CLIP)的发展,提示学习旨在提升模型迁移能力。传统方法依赖类别名称作为语义先验进行语义感知分类,但在实际中常仅有少量样本和类别实例信息,形成语义无关判别场景。文本到图像个性化方法通过学习新标记使模型生成未见概念,无需类别名称先验。本文首次揭示:文本反转所学的新标记同时具备生成与分类能力,若将每个类别视为单一概念。然而单概念文本反转在判别任务上表现不佳。为此,本文提出多类别文本反转(MC-TI),在标记更新中引入判别正则项,显著提升语义无关分类性能,同时保持生成能力。在12个覆盖多种场景的数据集上评估,结果表明MC-TI在分类与生成两方面均优于现有方法。

原文摘要 · Abstract (English)

With the advent of large pre-trained vision-language models such as CLIP, prompt learning methods aim to enhance the transferability of the CLIP model. They learn the prompt given few samples from the downstream task given the specific class names as prior knowledge, which we term as semantic-aware classification. However, in many realistic scenarios, we only have access to few samples and knowledge of the class names (e.g., when considering instances of classes). This challenging scenario represents the semantic-agnostic discriminative case. Text-to-Image (T2I) personalization methods aim to adapt T2I models to unseen concepts by learning new tokens and endowing these tokens with the capability of generating the learned concepts. These methods do not require knowledge of class names as a semantic-aware prior. Therefore, in this paper, we first explore Textual Inversion and reveal that the new concept tokens possess both generation and classification capabilities by regarding each category as a single concept. However, learning classifiers from single-concept textual inversion is limited since the learned tokens are suboptimal for the discriminative tasks. To mitigate this issue, we propose Multi-Class textual inversion, which includes a discriminative regularization term for the token updating process. Using this technique, our method MC-TI achieves stronger Semantic-Agnostic Classification while preserving the generation capability of these modifier tokens given only few samples per category. In the experiments, we extensively evaluate MC-TI on 12 datasets covering various scenarios, which demonstrates that MC-TI achieves superior results in terms of both classification and generation outcomes.

文本反转零样本分类生成模型少样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。