arXiv:2509.03494cs.CV2025-09

用像素级视觉提示高效适配大模型做无参考图像质量评估

Parameter-Efficient Adaptation of mPLUG-Owl2 via Pixel-Level Visual Prompts for NR-IQA

  • 仅训练60万参数(<0.01%),冻结原模型,通过像素空间提示适配
  • 在KADID-10k数据集上达到0.93的SRCC,媲美全量微调方法
  • 首次将像素级视觉提示用于无参考图像质量评估,适合低资源场景

本文提出一种新型参数高效的无参考图像质量评估(NR-IQA)方法,利用像素空间优化的视觉提示进行多模态大模型(MLLM)适配。与全量微调不同,该方法最多仅训练60万参数(占基础模型的<0.01%),并保持底层模型完全冻结。推理时,这些视觉提示通过加法与图像结合,输入mPLUG-Owl2模型,并以文本查询“评价图像的技术质量”进行处理。在包含合成、真实和AI生成失真类型的KADID-10k、KonIQ-10k及AGIQA-3k数据集上评估,性能可媲美全量微调方法与专用NR-IQA模型,在KADID-10k上达到0.93的SRCC。据我们所知,这是首个将像素空间视觉提示应用于NR-IQA的工作,实现了对低层视觉任务的高效MLLM适配。代码已公开于https://github.com/yahya-ben/mplug2-vp-for-nriqa。

原文摘要 · Abstract (English)

In this paper, we propose a novel parameter-efficient adaptation method for No- Reference Image Quality Assessment (NR-IQA) using visual prompts optimized in pixel-space. Unlike full fine-tuning of Multimodal Large Language Models (MLLMs), our approach trains only 600K parameters at most (< 0.01% of the base model), while keeping the underlying model fully frozen. During inference, these visual prompts are combined with images via addition and processed by mPLUG-Owl2 with the textual query "Rate the technical quality of the image." Evaluations across distortion types (synthetic, realistic, AI-generated) on KADID- 10k, KonIQ-10k, and AGIQA-3k demonstrate competitive performance against full finetuned methods and specialized NR-IQA models, achieving 0.93 SRCC on KADID-10k. To our knowledge, this is the first work to leverage pixel-space visual prompts for NR-IQA, enabling efficient MLLM adaptation for low-level vision tasks. The source code is publicly available at https: // github. com/ yahya-ben/ mplug2-vp-for-nriqa.

图像质量评估视觉提示参数高效多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。