对比人类感知,评估视觉语言模型在图像质量判断中的表现。
Vision-Language Models vs Human: Perceptual Image Quality Assessment
- 用心理物理实验数据作为基准,系统评测六种视觉语言模型。
- 色彩丰富度判断接近人类(相关系数ρ达0.93),但对比度判断较差。
- 模型对色彩的权重偏高,与人类倾向一致,且刺激差异越明显越可靠。
心理物理学实验仍是感知图像质量评估(IQA)最可靠的方法,但成本高、难以扩展,促使自动化方法的发展。本文研究视觉语言模型(VLMs)在对比度、色彩丰富度和整体偏好三个图像质量维度上是否能逼近人类感知判断。六种VLMs(四种专有模型,两种开源模型)与心理物理学数据进行对比。结果表明,模型表现具有显著属性依赖性:在色彩丰富度上对人类判断高度对齐(ρ最高达0.93),但在对比度上表现较弱;反之亦然。属性加权分析显示,大多数VLM在整体偏好判断中对色彩赋予更高权重,与心理物理学数据趋势一致。模型内部一致性分析揭示反直觉现象:最自洽的模型未必最贴近人类,说明响应变异性可能反映了对场景依赖感知线索的敏感性。此外,人类与模型的一致性随感知可分离性提高而增强,表明当刺激差异清晰时,模型更可靠。
原文摘要 · Abstract (English)
Psychophysical experiments remain the most reliable approach for perceptual image quality assessment (IQA), yet their cost and limited scalability encourage automated approaches. We investigate whether Vision Language Models (VLMs) can approximate human perceptual judgments across three image quality scales: contrast, colorfulness and overall preference. Six VLMs four proprietary and two openweight models are benchmarked against psychophysical data. This work presents a systematic benchmark of VLMs for perceptual IQA through comparison with human psychophysical data. The results reveal strong attribute dependent variability models with high human alignment for colorfulness (ρup to 0.93) underperform on contrast and vice-versa. Attribute weighting analysis further shows that most VLMs assign higher weights to colorfulness compared to contrast when evaluating overall preference similar to the psychophysical data. Intramodel consistency analysis reveals a counterintuitive tradeoff: the most self consistent models are not necessarily the most human aligned suggesting response variability reflects sensitivity to scene dependent perceptual cues. Furthermore, human-VLM agreement is increased with perceptual separability, indicating VLMs are more reliable when stimulus differences are clearly expressed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。