测试视觉语言模型对生成图像中显著缺陷的理解能力
SalArt-VQA: Diagnosing Whether VLMs Understand Salient Artifacts in Generated Images

- 设计细粒度诊断基准,评估模型对缺陷的定位与证据支持
- 20个模型中最强者仅在53.26%的图像上全答正确
- 高检测准确率下仍存在错误线索依赖和虚假缺陷指控
视觉语言模型(VLMs)被广泛用于检测生成图像中的可见伪影,但其对这些伪影的理解能力仍不清晰。单一图像级判断可能掩盖关键失败:模型虽正确标记伪影,却可能依赖错误视觉线索、定位错误区域或描述图像不存在的缺陷。为此,我们提出SalArt-VQA,一个针对生成图像中显著伪影理解的诊断基准。该基准包含950张图像和3,681道人类编写的多选题,涵盖伪影图像、匹配的真实参考图及配对的生成参考图。四个对齐问题类型评估伪影存在性检测、语义定位、空间定位和证据支持的缺陷识别,参考图分组测试校准与无缺陷时的弃权行为。在20个VLM中,最强模型在伪影图像上的检测召回率达99.37%,但仅在53.26%的图像上四类问题全部答对。对比伪影图像与无伪影参考图揭示敏感性与校准间的权衡:敏感模型常做出无证据的伪影声明,保守模型则通过漏检真实伪影来避免误报。结果表明,高检测准确率不等于基于局部视觉证据的真正理解。SalArt-VQA暴露了这些隐藏失败模式,提供了细粒度评估手段。
原文摘要 · Abstract (English)
Vision-language models (VLMs) are increasingly used to detect whether AI-generated images contain visible artifacts, yet their ability to analyze such artifacts remains poorly understood. A correct image-level decision can still hide important failures: a model may correctly flag an artifact while relying on the wrong visual cue, selecting the wrong region, or describing a defect that the image does not support. To evaluate these behaviors directly, we introduce SalArt-VQA, a diagnostic benchmark for fine-grained SALient ARTifact understanding in AI-generated images. SalArt-VQA contains 950 images and 3,681 human-authored multiple-choice questions spanning artifact images, matched real reference images, and paired generated reference images. Four aligned question types evaluate presence detection, semantic localization, spatial grounding, and evidence-grounded defect identification, while the reference splits test calibration and abstention when the annotated defect is absent. Across 20 VLMs, SalArt-VQA reveals failures that image-level detection accuracy hides: the strongest model reaches 99.37% detection recall on artifact images but answers all four artifact-side questions correctly on only 53.26% of images. Comparing artifact images with artifact-free references reveals a sensitivity-calibration tradeoff: sensitive models often make unsupported artifact claims, while conservative models avoid false alarms largely by missing real artifacts. These results show that high artifact detection accuracy alone does not imply grounded artifact understanding. SalArt-VQA exposes these hidden failure modes and provides a fine-grained evaluation of whether VLM artifact claims are supported by local visual evidence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。