arXiv:2507.01042cs.IRcs.AI2025-07被引 2

对比视觉语言模型跨领域表现,发现能力强者未必可靠。

Can Argus Judge Them All? Comparing VLMs Across Domains

  • 构建能力-可靠性双维度评估框架,看模型在不同数据集上是否稳定
  • Qwen-2.5VL-3B-Instruct能力最强,但CLIP延迟最低、内存占用最小
  • 适合选型时兼顾性能与稳定的工程师或系统设计者

视觉语言模型(VLMs)在检索、内容生成和决策支持等工业应用中日益普及,模型选择常依赖基准排名。这些排名主要基于检索、图像描述和推理任务,但相似任务表现的模型在不同数据集上行为差异显著,导致能力与可靠性之间存在差距。本文提出ARGUS-EVAL评估框架,从基准能力P(M)、跨数据集一致性CDC(M)、鲁棒性保留率RR(M)和效率E(M)四方面刻画模型行为。我们评估了CLIP、BLIP、LXMERT、Gemma-3-4B和Qwen-2.5VL-3B-Instruct在检索、描述和推理任务上的表现。结果表明,能力导向与可靠性导向的排名存在明显差异:Qwen-2.5VL-3B-Instruct在整体能力上最优(R@1 = 82.7%,BLEU-4 = 47.2%,CIDEr = 141.6,CDC = 0.91),而CLIP延迟最低(31 ms),内存占用最小(0.9 GB)。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) are increasingly used in industry VLM applications such as retrieval systems, content generation platforms, and decision-support workflows, where model selection is commonly guided by benchmark rankings. These rankings are largely determined by retrieval, captioning, and reasoning downstream tasks; however, models with similar task performance often show substantially different behavior across datasets. This creates a Capability-Reliability Gap between benchmark performance and observed model stability. We present ARGUS-EVAL, a capability-reliability-oriented evaluation framework for VLMs that characterizes model behavior through Benchmark Capability P(M), Cross-Dataset Consistency CDC(M), Robustness Retention RR(M), and Efficiency E(M). We evaluate CLIP, BLIP, LXMERT, Gemma-3-4B, and Qwen-2.5VL-3B-Instruct across retrieval, captioning, and reasoning downstream tasks. The results reveal notable differences between capability-oriented and reliability-oriented rankings. Qwen-2.5VL-3BInstruct achieves the strongest overall capability (R@1 = 82.7%, BLEU-4 = 47.2%, CIDEr = 141.6, CDC = 0.91), whereas CLIP records the lowest latency (31 ms) and memory footprint (0.9 GB).

视觉语言模型模型评估可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。