提出三类心理物理学测试,检验图像质量度量对人类低层视觉的模拟能力。
Evaluating quality metrics through the lenses of psychophysical measurements of low-level vision
- 基于人类视觉感知设计对比敏感度、掩蔽和匹配测试
- 34个度量中仅有LPIPS和MS-SSIM能较好预测对比掩蔽
- 揭示现有度量在高空间频率和对比恒常性上的缺陷
图像与视频质量度量(如SSIM、LPIPS、VMAF)旨在预测主观视觉质量,常被认为反映人眼感知原理。然而,多数度量依赖手工公式或数据驱动训练,缺乏对人类视觉模型的显式建模。本文引入一套全参考质量度量评估测试,用于检验其对低层人类视觉关键特性——对比敏感度、对比掩蔽和对比匹配的捕捉能力。我们对34个现有度量进行测试,发现LPIPS和MS-SSIM能较好预测对比掩蔽,而SSIM过度强调高空间频率,这一问题在MS-SSIM中得到缓解;此外,多数度量无法建模超阈值对比恒常性。结果表明,这些测试可揭示标准评估协议难以察觉的度量性质。
原文摘要 · Abstract (English)
Image and video quality metrics, such as SSIM, LPIPS, and VMAF, aim to predict perceived visual quality and are often assumed to reflect principles of human vision. However, relatively few metrics explicitly incorporate models of human perception, with most relying on hand-crafted formulas or data-driven training to approximate perceptual alignment. In this paper, we introduce a set of tests for full-reference quality metrics that evaluate their ability to capture key aspects of low-level human vision: contrast sensitivity, contrast masking, and contrast matching. These tests provide an additional framework for assessing both established and newly proposed metrics. We apply the tests to 34 existing quality metrics and highlight patterns in their behavior, including the ability of LPIPS and MS-SSIM to predict contrast masking and the tendency of SSIM to overemphasize high spatial frequencies, which is mitigated in MS-SSIM, and the general inability of metrics to model supra-threshold contrast constancy. Our results demonstrate how these tests can reveal properties of quality metrics that are not easily observed with standard evaluation protocols.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。