arXiv:2512.09573cs.CV2025-12被引 1

探究视觉语言模型对图像低层失真感知的缺陷与改进方法

Investigate the Low-level Visual Perception in Vision-Language based Image Quality Assessment

  • 设计低层失真分类任务,检验模型对模糊、噪声等基本失真的识别能力
  • 微调后识别准确率从14.92%提升至84.43%,显著改善失真感知
  • 强调视觉编码器对齐优化是实现可解释图像质量评估的关键

近年来,图像质量评估(IQA)利用多模态大语言模型(MLLMs)生成描述性解释。然而,尽管这些模型具备强大的视觉感知模块,却常无法可靠检测模糊、噪声和压缩等基本低层失真,且重复推理时评估结果不一致。这引发一个核心问题:基于MLLM的IQA系统是否真正感知了关键视觉特征?为此,我们引入一项低层失真感知任务,要求模型分类特定失真类型。组件级分析显示,尽管MLLM在结构上具备表示此类失真的能力,但容易过拟合训练模板,导致评分偏差。关键低层特征在视觉-语言对齐传递阶段被削弱或丢失。通过计算微调前后视觉特征与语义标记间的语义距离,我们发现提升视觉编码器对齐可显著增强失真识别精度,准确率从14.92%提高到84.43%。结果表明,在视觉编码器中加入专门约束,能强化文本可解释的视觉表征,使基于MLLM的流水线在以视觉为中心的任务中产生更连贯、可解释的推理。

原文摘要 · Abstract (English)

Recent advances in Image Quality Assessment (IQA) have leveraged Multi-modal Large Language Models (MLLMs) to generate descriptive explanations. However, despite their strong visual perception modules, these models often fail to reliably detect basic low-level distortions such as blur, noise, and compression, and may produce inconsistent evaluations across repeated inferences. This raises an essential question: do MLLM-based IQA systems truly perceive the visual features that matter? To examine this issue, we introduce a low-level distortion perception task that requires models to classify specific distortion types. Our component-wise analysis shows that although MLLMs are structurally capable of representing such distortions, they tend to overfit training templates, leading to biases in quality scoring. As a result, critical low-level features are weakened or lost during the vision-language alignment transfer stage. Furthermore, by computing the semantic distance between visual features and corresponding semantic tokens before and after component-wise fine-tuning, we show that improving the alignment of the vision encoder dramatically enhances distortion recognition accuracy, increasing it from 14.92% to 84.43%. Overall, these findings indicate that incorporating dedicated constraints on the vision encoder can strengthen text-explainable visual representations and enable MLLM-based pipelines to produce more coherent and interpretable reasoning in vision-centric tasks.

图像质量评估多模态模型视觉感知可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。