arXiv:2512.09555cs.CV2025-12中稿 · the ICONIP被引 2

让视觉语言模型像人一样推理图像质量,提升判断稳定性和可信度。

Building Reasonable Inference for Vision-Language Models in Blind Image Quality Assessment

  • 分两阶段训练:先学视觉特征,再仅凭特征推断质量
  • 在SPAQ/KONIQ上预测稳定性提升,相关性指标平均提高0.35
  • 适合追求可解释性与稳定性的图像质量评估研究者

近期盲图像质量评估(BIQA)进展主要依赖视觉语言模型(VLM),其语义推理能力暗示其可能以类人方式提取视觉特征、生成描述文本并推断质量。然而,这些模型常产生与最终评分矛盾的文本描述,且预测分数在推理过程中不稳定——这与人类推理不符。我们分析了导致矛盾评估和不稳定的因素:首先发现最终评分与生成的视觉特征之间关联较弱,逻辑连接不充分;其次,中间层解码显示模型频繁依赖少数候选词,加剧了预测波动。为此,我们提出一种两阶段微调方法,显式分离视觉感知与质量推断:第一阶段学习视觉特征,第二阶段仅基于这些特征进行质量判断。在SPAQ和KONIQ数据集上的实验表明,该方法将预测不稳定性从22.00%降至12.39%,并在LIVE、CSIQ、SPAQ、KONIQ上实现SRCC/PLCC平均提升0.3124/0.3507。进一步分析显示,该方法同时提升了推理过程的稳定性和可靠性。

原文摘要 · Abstract (English)

Recent progress in BIQA has been driven by VLMs, whose semantic reasoning abilities suggest that they might extract visual features, generate descriptive text, and infer quality in a human-like manner. However, these models often produce textual descriptions that contradict their final quality predictions, and the predicted scores can change unstably during inference - behaviors not aligned with human reasoning. To understand these issues, we analyze the factors that cause contradictory assessments and instability. We first estimate the relationship between the final quality predictions and the generated visual features, finding that the predictions are not fully grounded in the features and that the logical connection between them is weak. Moreover, decoding intermediate VLM layers shows that the model frequently relies on a limited set of candidate tokens, which contributes to prediction instability. To encourage more human-like reasoning, we introduce a two-stage tuning method that explicitly separates visual perception from quality inference. In the first stage, the model learns visual features; in the second, it infers quality solely from these features. Experiments on SPAQ and KONIQ demonstrate that our approach reduces prediction instability from 22.00% to 12.39% and achieves average gains of 0.3124/0.3507 in SRCC/PLCC across LIVE, CSIQ, SPAQ, and KONIQ compared to the baseline. Further analyses show that our method improves both stability and the reliability of the inference process.

图像质量评估视觉语言模型推理稳定性可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。