测试视觉语言模型的图像比较能力,发现多数模型无法保持对称性。
PairBench: Are Vision-Language Models Reliable at Comparing What They See?
- 用现成图像数据集构建可定制的对比评估框架
- 发现所有模型在对称性上表现差,且各有优劣
- 适合关注自动评估与排序任务的研究者
理解大型视觉语言模型(VLMs)在图像比较任务中的表现至关重要,但这一基础能力尚未得到系统评估。尽管VLMs被广泛用于需要比较判断的场景,如自动评价、重排序和检索增强生成,目前尚无系统性框架衡量其性能。本文提出PairBench,一个利用常见图像数据集评估VLMs作为可定制相似性工具的简单框架。该方法引入四项关键指标:与人类标注的一致性、配对顺序下的稳定性、分布平滑性以及提示控制能力。分析显示,无模型在所有指标上均表现优异,各具独特优势与缺陷。最令人担忧的是,多数模型无法维持对称的相似度评分。有趣的是,本基准上的表现与复杂任务常用基准高度相关,同时提供关于可控性、平滑性和顺序性的额外洞察。这使PairBench成为评估VLM在自动评价任务中表现的独特且全面的框架。
原文摘要 · Abstract (English)
Understanding how effectively large vision language models (VLMs) compare visual inputs is crucial across numerous applications, yet this fundamental capability remains insufficiently assessed. While VLMs are increasingly deployed for tasks requiring comparative judgment, including automated evaluation, re-ranking, and retrieval-augmented generation, no systematic framework exists to measure their performance in these scenarios. We present PairBench, a simple framework that evaluates VLMs as customizable similarity tools using widely available image datasets. Our approach introduces four key metrics for reliable comparison: alignment with human annotations, consistency across pair ordering, distribution smoothness, and controllability through prompting. Our analysis reveals that no model consistently excels across all metrics, with each demonstrating distinct strengths and weaknesses. Most concerning is the widespread inability of VLMs to maintain symmetric similarity scores. Interestingly, we demonstrate that performance on our benchmark strongly correlates with popular benchmarks used for more complex tasks, while providing additional metrics into controllability, smoothness and ordering. This makes PairBench a unique and comprehensive framework to evaluate the performance of VLMs for automatic evaluation depending on the task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。