用双流交互融合提升视觉Transformer在图像质量评估中的表现
Unleashing Vision Transformer Potential In Image Quality Assessment via Global-Local Adaptive Interaction

- 通过全局与局部特征交互,高效利用预训练Vision Transformer
- 在多个基准上实现更优预测精度,参数量显著减少
- 适合需要轻量化高精度图像质量评估的场景
在无参考图像质量评估(BIQA)领域,真实环境中多样复杂的失真使得准确预测感知质量仍具挑战。尽管现有方法取得较好精度,但其可扩展性受限于主观标注成本高及数据集规模小。近年来大规模预训练视觉模型虽具备强大语义表征能力,但在图像质量评估任务中受计算开销大、微调效率低制约。为此,本文提出全局-局部交互适配器(GLIA),采用双流特征提取与交互式全局-局部融合机制,有效利用预训练Vision Transformer。该方法同时保留全局语义信息与细粒度局部细节,在显著减少可训练参数的前提下,实现更高的预测精度与鲁棒性。多基准测试验证了该方法的有效性与优越性。
原文摘要 · Abstract (English)
In the field of Blind Image Quality Assessment (BIQA), accurately predicting the perceptual quality of authentically distorted images remains highly challenging due to the diverse and complex distortions present in natural environments. Although existing methods have achieved notable accuracy, their scalability is often constrained by the high cost of subjective annotation and the limited size of available datasets. Recent advances in large-scale pre-trained vision models have introduced powerful semantic and representational capabilities, yet their application to IQA tasks is hindered by substantial computational demands and suboptimal fine-tuning efficiency. To overcome these limitations, we introduce the Global-Local Interaction Adapter (GLIA), a novel framework that effectively harnesses pre-trained Vision Transformers through a dual-stream feature extraction mechanism coupled with interactive global-local fusion. By jointly retaining global semantic information and fine-grained local details, our approach delivers superior prediction accuracy and robustness while requiring significantly fewer trainable parameters. Extensive experiments on multiple benchmarks validate the effectiveness and superiority of our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。