arXiv:2601.20689cs.CVcs.AI2026-01被引 1

用小样本标注实现高质量图像评估,突破人工评分瓶颈。

Decoupling Perception and Calibration: Label-Efficient Image Quality Assessment Framework

  • 从大模型提取感知先验,轻量学生模型学习并校准质量判断
  • 仅需少量人类标注即可达到与全量标注相近的评估相关性
  • 适合标注资源有限但需高精度图像质量评估的场景

近期多模态大语言模型(MLLM)在图像质量评估(IQA)任务中表现强劲。然而,其适配过程计算成本高,仍依赖大量主观评分(MOS)。我们认为,基于MLLM的IQA核心瓶颈并非感知能力,而在于MOS尺度校准。为此提出LEAF框架:将MLLM教师模型的密集感知先验(包括逐点判断与成对偏好)蒸馏至轻量级学生回归器,并在少量MOS子集上进行校准,实现高效精准的主观评分对齐。在用户生成与AI生成图像的多个基准测试中,该方法显著降低对人工标注的需求,同时保持强MOS相关性,使轻量化IQA在标注预算受限时仍具可行性。

原文摘要 · Abstract (English)

Recent multimodal large language models (MLLMs) have demonstrated strong capabilities in image quality assessment (IQA) tasks. However, adapting such large-scale models is computationally expensive and still relies on substantial Mean Opinion Score (MOS) annotations. We argue that for MLLM-based IQA, the core bottleneck lies not in the quality perception capacity of MLLMs, but in MOS scale calibration. Therefore, we propose LEAF, a Label-Efficient Image Quality Assessment Framework that distills perceptual quality priors from an MLLM teacher into a lightweight student regressor, enabling MOS calibration with minimal human supervision. Specifically, the teacher conducts dense supervision through point-wise judgments and pair-wise preferences, with an estimate of decision reliability. Guided by these signals, the student learns the teacher's quality perception patterns through joint distillation and is calibrated on a small MOS subset to align with human annotations. Experiments on both user-generated and AI-generated IQA benchmarks demonstrate that our method significantly reduces the need for human annotations while maintaining strong MOS-aligned correlations, making lightweight IQA practical under limited annotation budgets.

图像质量评估少样本学习知识蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。