用轻量适配提升视觉语言模型在盲图像质量评估中的表现
Parameter-Efficient Adaptation of a Multi-Stream Vision-Language Framework for Blind Image Quality Assessment

- 融合自然场景统计与冻结的SigLIP/CLIP特征,用低秩适配微调骨干网络
- 仅训练0.23%参数,使模型在真实测试中性能提升最高达+0.357 SROCC
- 揭示图像级划分导致误判,参考级评估更能反映真实泛化能力
盲图像质量评估(BIQA)在无原始参考图像的情况下预测感知质量,是图像压缩、传输和修复等应用的基础。近年来,越来越多的BIQA方法依赖大型视觉语言模型(VLM)。尽管冻结的VLM可避免昂贵的全量微调,但其性能损失程度尚不明确,且何时真正需要适配骨干网络仍不清楚。这一问题因合成失真基准中普遍采用图像级划分而复杂化——同一参考图像的失真版本可能同时出现在训练和测试集中,造成内容重叠,人为抬高冻结表示的性能,掩盖其真实泛化能力,可能导致对骨干适配价值的错误判断。为此,本文提出一种高效框架,将自然场景统计描述符与冻结的SigLIP和CLIP-H嵌入通过轻量回归头融合,并使用参数高效的低秩适配(LoRA)对SigLIP骨干进行微调,仅训练其0.23%的参数。在六大数据集上,对比图像级与参考级协议下的冻结与适配模型,发现图像级划分使冻结特征的SROCC被夸大高达0.44,掩盖了真实难度的显著差异;而参考级评估揭示出:当冻结特征泛化能力弱时,LoRA适配带来最大收益(如在TID2013上提升+0.357 SROCC),而在原已强大的情况下则增益甚微。
原文摘要 · Abstract (English)
Blind image quality assessment (BIQA) predicts perceived image quality without access to a pristine reference and is fundamental to applications such as image compression, transmission, and restoration. Recent BIQA methods increasingly rely on large vision-language models (VLMs). Although frozen VLMs provide an efficient alternative to computationally expensive full fine-tuning, it remains unclear how much performance is sacrificed by not adapting the backbone and, more importantly, under what conditions such adaptation is truly beneficial. Answering this question, however, is complicated by the widespread use of image-level splitting on synthetic-distortion benchmarks, where distorted versions of the same reference image can appear in both training and test partitions. This content overlap artificially inflates the apparent performance of frozen representations, masking their true generalization ability and potentially leading to incorrect conclusions about the value of backbone adaptation. We therefore address these two issues jointly. We develop an efficient BIQA framework that fuses a natural-scene-statistics descriptor with frozen SigLIP and CLIP-H embeddings through a lightweight regression head, and then apply parameter-efficient Low-Rank Adaptation (LoRA) to the SigLIP backbone, training only $0.23\%$ of its parameters. Evaluating both frozen and adapted models across six datasets under image-level and reference-level protocols, we find that image-level splitting inflates frozen-feature SROCC by up to $0.44$ and masks wide variation in true difficulty, which reference-level evaluation reveals. Under this content-independent protocol, LoRA adaptation recovers performance in proportion to the exposed difficulty, with the largest gains where frozen features generalize poorly (up to $+0.357$ SROCC on TID2013) and little benefit where they are already strong.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。