arXiv:2411.12791cs.CVeess.IV2024-11AAAI被引 7

不重训模型,用降质图像修正大模型对画质判断的偏见。

Mitigating Perception Bias: A Training-Free Approach to Enhance LMM for Image Quality Assessment

  • 通过保持语义不变的降质操作生成弱化画质的图像对。
  • 在多个数据集上使大模型画质评估准确率显著提升。
  • 适合想快速提升现成大模型画质评估能力的研究者使用。

尽管大型多模态模型(LMM)在高层视觉任务中表现优异,其在图像质量评估(IQA)上的能力仍受限。主要原因是这些模型主要针对高层任务(如图像描述)训练,强调在不同质量下提取统一的图像语义,导致其感知存在语义敏感而质量不敏感的偏差,迫使模型在评分时过度依赖语义信息。本文提出一种无需训练的去偏框架,通过削弱语义带来的干扰来修正图像质量预测。具体而言,我们探索了几种语义保持但严重降低画质的失真方式,将这些失真应用于查询或测试图像,使降质后图像仍可识别语义但被判定为低质量。推理时,同时输入原图及其降质版本,并提示模型‘若降质图被判为差,则原图质量应如何评判’。该先验条件使模型对所有降质图像一致判为差,从而校准其质量感知。最终,利用条件概率模型聚合不同降质版本下的评分结果。在多个IQA数据集上的实验表明,该框架能持续提升LMM性能。

原文摘要 · Abstract (English)

Despite the impressive performance of large multimodal models (LMMs) in high-level visual tasks, their capacity for image quality assessment (IQA) remains limited. One main reason is that LMMs are primarily trained for high-level tasks (e.g., image captioning), emphasizing unified image semantics extraction under varied quality. Such semantic-aware yet quality-insensitive perception bias inevitably leads to a heavy reliance on image semantics when those LMMs are forced for quality rating. In this paper, instead of retraining or tuning an LMM costly, we propose a training-free debiasing framework, in which the image quality prediction is rectified by mitigating the bias caused by image semantics. Specifically, we first explore several semantic-preserving distortions that can significantly degrade image quality while maintaining identifiable semantics. By applying these specific distortions to the query or test images, we ensure that the degraded images are recognized as poor quality while their semantics mainly remain. During quality inference, both a query image and its corresponding degraded version are fed to the LMM along with a prompt indicating that the query image quality should be inferred under the condition that the degraded one is deemed poor quality. This prior condition effectively aligns the LMM's quality perception, as all degraded images are consistently rated as poor quality, regardless of their semantic variance. Finally, the quality scores of the query image inferred under different prior conditions (degraded versions) are aggregated using a conditional probability model. Extensive experiments on various IQA datasets show that our debiasing framework could consistently enhance the LMM performance.

图像质量多模态去偏无训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。