arXiv:2510.22571cs.CVcs.AI2025-10被引 3

首个系统评估视觉语言模型物体状态理解能力的基准测试

STATUS Bench: A Rigorous Benchmark for Evaluating Object State Understanding in Vision-Language Models

  • 设计三任务联合评估框架,同时检验状态识别、图像检索与状态变化判断
  • 实测多数开源模型零样本性能仅达随机水平,微调后效果显著提升
  • 提供1300万条半自动生成描述数据集,推动该领域研究发展

物体状态识别旨在判断物体的具体状态,如位置状态(开/关)和功能状态(开启/关闭)。尽管近期视觉语言模型(VLMs)能完成多种多模态任务,但其对物体状态细微差异的理解精度仍不明确。为此,我们提出首个严谨评估VLM物体状态理解能力的基准——STATUS Bench。该基准采用新颖的评估范式,要求模型同时完成三项任务:物体状态识别(OSI)、图像检索(IR)和状态变化识别(SCI)。所有任务基于我们人工精心构建的数据集,包含图像对及其对应的物体状态与状态变化描述。此外,我们构建了大规模训练数据集STATUS Train,包含1300万条半自动生成的描述,是当前该领域最大资源。实验表明,STATUS Bench可实现严格的一致性评估,揭示当前主流VLM仍难以捕捉细微状态差异。令人意外的是,在新评估框架下,多数开源模型零样本表现接近随机水平。经在STATUS Train上微调后,Qwen2.5-VL性能已可媲美Gemini 2.0 Flash。这些发现凸显了STATUS Bench与Train在推进物体状态识别研究中的必要性。

原文摘要 · Abstract (English)

Object state recognition aims to identify the specific condition of objects, such as their positional states (e.g., open or closed) and functional states (e.g., on or off). While recent Vision-Language Models (VLMs) are capable of performing a variety of multimodal tasks, it remains unclear how precisely they can identify object states. To alleviate this issue, we introduce the STAte and Transition UnderStanding Benchmark (STATUS Bench), the first benchmark for rigorously evaluating the ability of VLMs to understand subtle variations in object states in diverse situations. Specifically, STATUS Bench introduces a novel evaluation scheme that requires VLMs to perform three tasks simultaneously: object state identification (OSI), image retrieval (IR), and state change identification (SCI). These tasks are defined over our fully hand-crafted dataset involving image pairs, their corresponding object state descriptions and state change descriptions. Furthermore, we introduce a large-scale training dataset, namely STATUS Train, which consists of 13 million semi-automatically created descriptions. This dataset serves as the largest resource to facilitate further research in this area. In our experiments, we demonstrate that STATUS Bench enables rigorous consistency evaluation and reveal that current state-of-the-art VLMs still significantly struggle to capture subtle object state distinctions. Surprisingly, under the proposed rigorous evaluation scheme, most open-weight VLMs exhibited chance-level zero-shot performance. After fine-tuning on STATUS Train, Qwen2.5-VL achieved performance comparable to Gemini 2.0 Flash. These findings underscore the necessity of STATUS Bench and Train for advancing object state recognition in VLM research.

视觉语言模型状态识别基准测试多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。