arXiv:2604.12813cs.CVcs.MM2026-04

用轻量校准分支让大模型快速适配新视频质量评估场景

DPC-VQA: Decoupling Quality Perception and Residual Calibration for Video Quality Assessment

论文配图:DPC-VQA: Decoupling Quality Perception and Residual Calibration for Video Quality Assessment
图 1 · 摘自论文原文
  • 冻结大模型提取感知先验,新增轻量分支校准质量评分
  • 仅需20%标注数据和不到2%参数量,性能媲美主流方法
  • 适合缺乏标注数据的视频质量评估新场景

近期多模态大语言模型在视频质量评估任务中表现优异,但适应新场景仍需大量重训练和昂贵的平均意见分(MOS)标注。本文认为,预训练的多模态大语言模型已具备有用的感知先验,主要挑战在于高效校准该先验至目标MOS空间。基于此,提出DPC-VQA框架,将感知与校准解耦:使用冻结的多模态大语言模型提供基础质量估计和感知先验,通过轻量校准分支预测目标场景的残差修正。该设计避免了昂贵的端到端重训练,同时保持可靠性能,显著降低训练和数据成本。在用户生成内容(UGC)和AI生成内容(AIGC)基准上的大量实验表明,DPC-VQA性能优于代表性基线,且仅使用常规多模态大语言模型方法不到2%的可训练参数,以及仅20%的MOS标签仍保持有效。代码将在发表后开源。

原文摘要 · Abstract (English)

Recent multimodal large language models (MLLMs) have shown promising performance on video quality assessment (VQA) tasks. However, adapting them to new scenarios remains expensive due to large-scale retraining and costly mean opinion score (MOS) annotations. In this paper, we argue that a pretrained MLLM already provides a useful perceptual prior for VQA, and that the main challenge is to efficiently calibrate this prior to the target MOS space. Based on this insight, we propose DPC-VQA, a decoupling perception and calibration framework for video quality assessment. Specifically, DPC-VQA uses a frozen MLLM to provide a base quality estimate and perceptual prior, and employs a lightweight calibration branch to predict a residual correction for target-scenario adaptation. This design avoids costly end-to-end retraining while maintaining reliable performance with lower training and data costs. Extensive experiments on both user-generated content (UGC) and AI-generated content (AIGC) benchmarks show that DPC-VQA achieves competitive performance against representative baselines, while using less than 2% of the trainable parameters of conventional MLLM-based VQA methods and remaining effective with only 20% of MOS labels. The code will be released upon publication.

视频质量评估大模型轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。