arXiv:2410.19808cs.CVcs.AI2024-10被引 2

评测视觉语言模型的定位能力,发现最强模型仍比人类差10%以上

LocateBench: Evaluating the Locating Ability of Vision Language Models

  • 构建专用于评估视觉定位能力的高质量基准数据集
  • 测试多类提示方法,发现最强模型GPT-4o准确率仍低于人类10%以上
  • 为提升模型空间理解能力提供可复现的评估标准,适合研究视觉定位的学者

根据自然语言指令在图像中定位物体的能力对众多实际应用至关重要。本文提出LocateBench,一个专注于评估该能力的高质量基准。我们实验了多种提示策略,并测量多个大型视觉语言模型的定位准确率。结果显示,即使是最强的模型GPT-4o,其准确率也比人类低超过10%。

原文摘要 · Abstract (English)

The ability to locate an object in an image according to natural language instructions is crucial for many real-world applications. In this work we propose LocateBench, a high-quality benchmark dedicated to evaluating this ability. We experiment with multiple prompting approaches, and measure the accuracy of several large vision language models. We find that even the accuracy of the strongest model, GPT-4o, lags behind human accuracy by more than 10%.

视觉定位模型评估多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。