arXiv:2606.04282cs.CV2026-06

首个评估多模态大模型定位能力的标准化基准,覆盖四类核心任务。

FindIt: A Format-Informed Visual Detection Benchmark for Generalist Multimodal LLMs

论文配图:FindIt: A Format-Informed Visual Detection Benchmark for Generalist Multimodal LLMs
图 1 · 摘自论文原文
  • 构建统一框架,规范输入输出格式与评估流程。
  • 发现当前模型对格式要求敏感,微小变化即导致失败。
  • 适合研究多模态模型定位能力与系统集成的开发者。

多模态大语言模型(MLLMs)主要在自由形式的视觉-语言任务(如视觉问答、图像描述)上进行评估。然而,其实际应用正迅速扩展到更结构化的计算机视觉场景,用户常通过提示模型执行以定位为中心的任务(如目标检测),尤其是在更大的代理或决策系统中。尽管如此,目前尚无标准化基准能大规模系统评估此类能力。本文提出首个专为评估通用多模态大模型提示式定位能力而设计的综合性基准。该基准涵盖四大核心任务类别:目标检测、指代表达检测、实例级检测和基于视频的检测。为实现一致且公平的评估,我们开发了统一框架,标准化输入、强制可解析的边界框输出,并定义透明的评估协议。利用该套件,我们评估了多种开源与专有MLLMs,深入分析其性能与局限。除准确率外,还考察模型对输出格式规范的遵循能力,结果显示当前系统对格式约束极为敏感,即使在微小变化下也难以泛化。结果揭示了现有顶尖模型在定位场景中的优势与不足,为改进多模态模型设计与评估指明方向。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) are predominantly evaluated on free-form vision-language tasks such as visual question answering, captioning, and summarization. However, their practical use is rapidly expanding to more structured computer vision settings, where users prompt models to perform localization-centric tasks such as object detection, often within larger agentic or decision-making systems. Despite this shift, there is currently no standardized benchmark that systematically evaluates these capabilities at scale. In this work, we introduce the first comprehensive benchmark specifically designed to assess the promptable localization abilities of generalist MLLMs. Our benchmark spans four core task categories: object detection, referring expression detection, instance-level detection, and video-based detection. To enable consistent and fair evaluation, we develop a unified framework that standardizes inputs, enforces parsable bounding box outputs, and defines transparent evaluation protocols across tasks. Using this suite, we evaluate a diverse set of open-source and proprietary MLLMs, providing an in-depth analysis of their performance and limitations. Beyond accuracy, we examine models' ability to adhere to output format specifications, showing that current systems are highly sensitive to formatting constraints and often fail to generalize even to minor variations. Our results highlight both the strengths and shortcomings of state-of-the-art MLLMs in localization settings, and point toward important directions for improving multimodal model design and evaluation.

多模态定位检测基准测试大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。