用检索增强生成提升大模型图像质量评估能力,无需训练。
Enhancing Image Quality Assessment Ability of LMMs via Retrieval-Augmented Generation
- 通过检索相似但质量不同的参考图构建提示,引导模型判断
- 在多个数据集上显著提升大模型的图像质量评估性能
- 无需微调,适合资源受限场景下的高质量评估需求
大型多模态模型(LMMs)在低层视觉感知任务中展现出巨大潜力,尤其在图像质量评估(IQA)方面表现出强大的零样本能力。然而,达到顶尖性能通常需要计算成本高昂的微调方法,旨在对齐输出中与质量相关的标记分布与图像质量等级。受近期无训练工作启发,我们提出IQARAG——一种新颖的无训练框架,用于增强LMMs的IQA能力。IQARAG利用检索增强生成(RAG)技术,为输入图像检索若干语义相似但质量各异的参考图像及其对应的平均意见分数(MOS)。这些检索图像与输入图像共同构成特定提示,为LMM提供视觉感知锚点。IQARAG包含三个关键阶段:检索特征提取、图像检索、集成与质量评分生成。在多个多样化的IQA数据集(包括KADID、KonIQ、LIVE Challenge和SPAQ)上的广泛实验表明,所提出的IQARAG能有效提升LMMs的IQA性能,为质量评估提供一种资源高效的替代方案。
原文摘要 · Abstract (English)
Large Multimodal Models (LMMs) have recently shown remarkable promise in low-level visual perception tasks, particularly in Image Quality Assessment (IQA), demonstrating strong zero-shot capability. However, achieving state-of-the-art performance often requires computationally expensive fine-tuning methods, which aim to align the distribution of quality-related token in output with image quality levels. Inspired by recent training-free works for LMM, we introduce IQARAG, a novel, training-free framework that enhances LMMs' IQA ability. IQARAG leverages Retrieval-Augmented Generation (RAG) to retrieve some semantically similar but quality-variant reference images with corresponding Mean Opinion Scores (MOSs) for input image. These retrieved images and input image are integrated into a specific prompt. Retrieved images provide the LMM with a visual perception anchor for IQA task. IQARAG contains three key phases: Retrieval Feature Extraction, Image Retrieval, and Integration & Quality Score Generation. Extensive experiments across multiple diverse IQA datasets, including KADID, KonIQ, LIVE Challenge, and SPAQ, demonstrate that the proposed IQARAG effectively boosts the IQA performance of LMMs, offering a resource-efficient alternative to fine-tuning for quality assessment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。