arXiv:2412.13947cs.CV2024-12被引 2

让CLIP凭描述识别物体部件,突破仅靠类别名的局限

Real Classification by Description: Extending CLIP's Limits of Part Attributes Recognition

  • 用无类别名的属性描述训练,测试模型真实理解能力
  • 在6个基准上提升细粒度属性识别性能,最高增益达12.7%
  • 新架构融合多分辨率特征,适合视觉语言模型研究者

本研究定义并挑战零样本‘真实’描述分类任务,评估视觉语言模型(如CLIP)仅依据描述性属性而非对象类别名称进行分类的能力。该方法揭示了当前模型在理解复杂物体描述上的局限,推动其超越简单识别。为此,我们引入新挑战,并发布六个主流细粒度基准的描述数据集,去除对象名称以促进真正的零样本学习。同时,提出通过ImageNet21k多样类别与大语言模型生成的丰富属性描述对CLIP进行定向训练,提升其属性检测能力。此外,设计一种改进的CLIP架构,利用多分辨率信息增强细粒度部件属性识别。实验表明,该方法显著提升CLIP在六个主流基准及PACO数据集上的表现。代码已开源。

原文摘要 · Abstract (English)

In this study, we define and tackle zero shot "real" classification by description, a novel task that evaluates the ability of Vision-Language Models (VLMs) like CLIP to classify objects based solely on descriptive attributes, excluding object class names. This approach highlights the current limitations of VLMs in understanding intricate object descriptions, pushing these models beyond mere object recognition. To facilitate this exploration, we introduce a new challenge and release description data for six popular fine-grained benchmarks, which omit object names to encourage genuine zero-shot learning within the research community. Additionally, we propose a method to enhance CLIP's attribute detection capabilities through targeted training using ImageNet21k's diverse object categories, paired with rich attribute descriptions generated by large language models. Furthermore, we introduce a modified CLIP architecture that leverages multiple resolutions to improve the detection of fine-grained part attributes. Through these efforts, we broaden the understanding of part-attribute recognition in CLIP, improving its performance in fine-grained classification tasks across six popular benchmarks, as well as in the PACO dataset, a widely used benchmark for object-attribute recognition. Code is available at: https://github.com/ethanbar11/grounding_ge_public.

视觉语言模型细粒度识别CLIP改进

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。