提出新指标,评估视觉描述符的质量与模型预训练数据的匹配度。
Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor
- 用对齐度和相似性指标替代准确率,衡量描述符质量。
- 发现描述符在预训练数据中的出现频率影响其表现。
- 适合研究视觉语言模型描述符生成与对齐机制的研究者。
基于文本的视觉描述符(从简单类别名到描述性短语)广泛用于视觉概念发现和图像分类任务中,其有效性取决于语义清晰度、在视觉语言模型(VLM)预训练数据中的覆盖程度以及作为有意义表征空间的能力。本文系统分析描述符质量的两个关键维度:(1) 表示能力,(2) 与 VLM 预训练数据的关系。评估了从零样本大模型生成提示到迭代优化描述符的一系列生成方法。受表示对齐和语言理解思想启发,提出两种基于对齐的度量——全局对齐(Global Alignment)与 CLIP 相似性(CLIP Similarity),突破传统准确率局限。这些指标揭示不同描述符生成策略如何与基础模型特性相互作用,为超越准确率评估提供了新视角。
原文摘要 · Abstract (English)
Text-based visual descriptors--ranging from simple class names to more descriptive phrases--are widely used in visual concept discovery and image classification with vision-language models (VLMs). Their effectiveness, however, depends on a complex interplay of factors, including semantic clarity, presence in the VLM's pre-training data, and how well the descriptors serve as a meaningful representation space. In this work, we systematically analyze descriptor quality along two key dimensions: (1) representational capacity, and (2) relationship with VLM pre-training data. We evaluate a spectrum of descriptor generation methods, from zero-shot LLM-generated prompts to iteratively refined descriptors. Motivated by ideas from representation alignment and language understanding, we introduce two alignment-based metrics--Global Alignment and CLIP Similarity--that move beyond accuracy. These metrics shed light on how different descriptor generation strategies interact with foundation model properties, offering new ways to study descriptor effectiveness beyond accuracy evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。