测试视觉语言模型对图像失真类型的识别能力,发现其表现远不如人类。
DistortBench: Benchmarking Vision Language Models on Image Distortion Identification

- 构建13,500题多选基准,覆盖27种失真类型与五级严重程度。
- 最佳模型准确率仅61.9%,低于人类多数投票基准(65.7%)。
- 揭示模型规模与性能无单调关系,且不同模型家族响应模式各异。
视觉语言模型(VLMs)在内容审核、图像修复和质量监控等场景中日益重要,但其对低层图像退化的感知能力仍不清晰。本文提出DistortBench,一个针对无参考失真感知的诊断性基准。该基准包含13,500道四选一问题,涵盖27种失真类型、六类感知类别和五级严重程度:其中25种失真沿用KADID-10k校准,新增两种旋转失真采用单调角度分级。我们评估了18个VLM,包括来自五个系列的17个开源模型及1个专有模型。尽管在高层视觉语言任务中表现优异,最佳模型准确率仅为61.9%,略低于人类多数投票基准(65.7%),个体平均为60.2%,表明当前VLM在低层感知理解方面仍存在显著短板。分析进一步揭示模型规模与性能间缺乏单调增长,多数基础-思维配对性能下降,且各模型家族对严重程度的响应模式各异。我们希望DistortBench能成为衡量并提升VLM低层视觉感知能力的重要工具。
原文摘要 · Abstract (English)
Vision-language models (VLMs) are increasingly used in settings where sensitivity to low-level image degradations matters, including content moderation, image restoration, and quality monitoring. Yet their ability to recognize distortion type and severity remains poorly understood. We present DistortBench, a diagnostic benchmark for no-reference distortion perception in VLMs. DistortBench contains 13,500 four-choice questions covering 27 distortion types, six perceptual categories, and five severity levels: 25 distortions inherit KADID-10k calibrations, while two added rotation distortions use monotonic angle-based levels. We evaluate 18 VLMs, including 17 open-weight models from five families and one proprietary model. Despite strong performance on high-level vision-language tasks, the best model reaches only 61.9% accuracy, just below the human majority-vote baseline of 65.7% (average individual: 60.2%), indicating that low-level perceptual understanding remains a major weakness of current VLMs. Our analysis further reveals weak and non-monotonic scaling with model size, performance drops in most base--thinking pairs, and distinct severity-response patterns across model families. We hope DistortBench will serve as a useful benchmark for measuring and improving low-level visual perception in VLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。