arXiv:2508.06092cs.CV2025-08被引 3

用视觉语言模型提升视频质量评估,仅需少量训练参数。

Q-CLIP: Unleashing the Power of Vision-Language Models for Video Quality Assessment through Unified Cross-Modal Adaptation

  • 通过共享跨模态适配器增强视觉与文本表征,仅训练少量参数。
  • 引入可学习的质量提示词,提升对细微画质差异的感知能力。
  • 适合追求高效高精度视频质量评估的研究者和工程师。

准确高效的视频质量评估(VQA)一直是研究难点。现有主流方法通常在大规模分类数据集(如ImageNet、Kinetics-400)上预训练,再在VQA数据集上微调,但面临两大挑战:(1) 仅迁移语义知识不足以覆盖视频质量多维度因素(如语义、失真、运动、美学);(2) 大规模预训练计算开销巨大,常为直接训练VQA数据集的数十甚至上百倍。近期,视觉语言模型(VLMs)在多种视觉任务中展现出强大泛化能力,并初显在质量评估中的潜力。本文提出Q-CLIP,首个完全基于VLM的VQA框架。Q-CLIP通过一个仅含极少可训练参数的共享跨模态适配器(SCMA)增强视觉与文本表示,显著降低计算成本。同时,引入五组可学习的质量等级提示词,引导模型捕捉细微画质变化,进一步提升敏感度。此外,研究不同帧采样策略影响,发现基于帧差的采样在跨数据集上具有更好泛化性能。大量实验表明,Q-CLIP在多个VQA数据集上表现优异。

原文摘要 · Abstract (English)

Accurate and efficient Video Quality Assessment (VQA) has long been a key research challenge. Current mainstream VQA methods typically improve performance by pretraining on large-scale classification datasets (e.g., ImageNet, Kinetics-400), followed by fine-tuning on VQA datasets. However, this strategy presents two significant challenges: (1) merely transferring semantic knowledge learned from pretraining is insufficient for VQA, as video quality depends on multiple factors (e.g., semantics, distortion, motion, aesthetics); (2) pretraining on large-scale datasets demands enormous computational resources, often dozens or even hundreds of times greater than training directly on VQA datasets. Recently, Vision-Language Models (VLMs) have shown remarkable generalization capabilities across a wide range of visual tasks, and have begun to demonstrate promising potential in quality assessment. In this work, we propose Q-CLIP, the first fully VLMs-based framework for VQA. Q-CLIP enhances both visual and textual representations through a Shared Cross-Modal Adapter (SCMA), which contains only a minimal number of trainable parameters and is the only component that requires training. This design significantly reduces computational cost. In addition, we introduce a set of five learnable quality-level prompts to guide the VLMs in perceiving subtle quality variations, thereby further enhancing the model's sensitivity to video quality. Furthermore, we investigate the impact of different frame sampling strategies on VQA performance, and find that frame-difference-based sampling leads to better generalization performance across datasets. Extensive experiments demonstrate that Q-CLIP exhibits excellent performance on several VQA datasets.

视频质量视觉语言模型跨模态适配提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。