arXiv:2608.07435cs.CVcs.AI2026-08

自动化构建视觉语言模型压力测试,发现模型依赖先验而非真实图像证据。

SABRE: Scalable and Automated Benchmarking of VLMs under Stress

论文配图:SABRE: Scalable and Automated Benchmarking of VLMs under Stress
图 1 · 摘自论文原文
  • 将任务设计自动转为带标注的图像与问题对,支持大规模生成。
  • 6个模型平均准确率仅22.6%,暴露出对图像证据的严重依赖。
  • 适合评估模型鲁棒性,尤其关注视觉推理可靠性研究者使用。

视觉语言模型(VLMs)发展迅速,但基准测试滞后,难以发现其弱点。构建压力测试成本高:样本需满足控制条件、可回答且能挑战现有模型。我们提出SABRE,一个可扩展的自动化流水线,将测试设计模板(Test Primer)转化为结构化规范、生成或编辑后的图像及问答对。通过过滤模型剔除已解决样本,人工审核验证有效性并修正标注与局部图像修复。我们以SABRE-Prior为例,检验VLM是否遵循视觉证据而非依赖世界先验——对常见物体与场景的预设认知。该数据集包含600张图像和1000个问题,涵盖情境(熟悉场景中的异常实体)、纹理(反事实材质)、属性(非标准组件数量)和语言诱导(语言暗示但图像不支持的答案)。在六个VLM上,宏观平均准确率介于17.8%至31.3%之间(均值22.6%)。真实图像属性控制组对过滤模型同样具有挑战性。SABRE-Counting和SABRE-Spatial试点证明该流程可拓展至其他压力测试场景。这些结果确立SABRE为可重复使用的VLM压力测试构建与更新框架,而非单一固定基准。

原文摘要 · Abstract (English)

Vision-language models (VLMs) are improving rapidly, but benchmark development lags behind, making weaknesses hard to identify. Building stress tests is costly: samples must satisfy controlled conditions, remain answerable, and challenge current models. We present SABRE, a scalable, automated pipeline that converts a Test Primer (a Markdown Task Design with Data Schema) into structured specifications, generated or edited images, and question-answer pairs. Automated filtering removes candidates solved by a Filtering VLM, while human review verifies candidate validity and supports annotation correction and localized image repair. We instantiate SABRE-Prior to test whether VLMs follow visual evidence instead of relying on world priors -- learned expectations about familiar objects and scenes. Its 600 images and 1,000 questions span Context (unexpected entities in familiar scenes), Texture (counterfactual materials), Attribute (noncanonical component counts), and Language Elicitation (answers suggested by language but unsupported by the image). Across six VLMs, macro-average accuracy ranges from 17.8% to 31.3% (22.6% mean). A real-image Attribute control is comparably difficult for the Filtering VLM. SABRE-Counting and SABRE-Spatial pilots show that the workflow supports other stress-test settings. These results establish SABRE as a reusable framework for constructing and refreshing VLM stress tests rather than a single fixed benchmark.

视觉语言模型压力测试鲁棒性评估自动化构建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。