arXiv:2511.07812cs.CV2025-11AAAI被引 5

改进多模态大模型的图像质量评估,解决评分不准问题

Revisiting MLLM Based Image Quality Assessment: Errors and Remedy

  • 引入轻量回归模块与专用评分标记,修正离散输出与连续评分的不匹配
  • 在多个图像质量评估基准上达到当前最优性能,跨数据集泛化能力强
  • 适合需要高精度图像质量判断的研究者和工业应用

多模态大语言模型(MLLM)的快速发展推动了图像质量评估(IQA)任务。然而,MLLM的离散标记输出与IQA所需的连续评分之间存在本质不匹配,严重制约了其性能。此前方法将离散标记转换为连续分数常引入误差,且“良好”等语义标记造成语义混淆,削弱了模型在相关任务中的能力。本文对错误根源进行理论分析,并提出简单有效的框架Q-Scorer:在MLLM流程中加入轻量回归模块与专用于IQA的评分标记。大量实验表明,Q-Scorer在多个IQA基准上达到领先性能,对混合数据集有良好泛化能力,且可与其他方法协同提升效果。

原文摘要 · Abstract (English)

The rapid progress of multi-modal large language models (MLLMs) has boosted the task of image quality assessment (IQA). However, a key challenge arises from the inherent mismatch between the discrete token outputs of MLLMs and the continuous nature of quality scores required by IQA tasks. This discrepancy significantly hinders the performance of MLLM-based IQA methods. Previous approaches that convert discrete token predictions into continuous scores often suffer from conversion errors. Moreover, the semantic confusion introduced by level tokens (e.g., ``good'') further constrains the performance of MLLMs on IQA tasks and degrades their original capabilities for related tasks. To tackle these problems, we provide a theoretical analysis of the errors inherent in previous approaches and, motivated by this analysis, propose a simple yet effective framework, Q-Scorer. This framework incorporates a lightweight regression module and IQA-specific score tokens into the MLLM pipeline. Extensive experiments demonstrate that Q-Scorer achieves state-of-the-art performance across multiple IQA benchmarks, generalizes well to mixed datasets, and further improves when combined with other methods.

图像质量评估多模态模型评分预测大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。