arXiv:2511.07290eess.IVcs.CV2025-11被引 5

用视频描述辅助评估压缩视频质量,无需人工标注。

CAMP-VQA: Caption-Embedded Multimodal Perception for No-Reference Quality Assessment of Compressed Video

  • 利用视觉语言模型生成细粒度质量描述,融合视频元数据与帧间变化片段。
  • 在多个UGC数据集上达到SRCC 0.928、PLCC 0.938的领先性能。
  • 适合视频平台自动质量评估,无需昂贵的人工标注。

YouTube、TikTok等平台上的用户生成内容(UGC)普及,使得无参考(NR)视频质量评估(VQA)对优化视频分发至关重要。然而,非专业拍摄及平台转码带来的复杂特征,使现有NR-VQA模型在推断主观评分时仍受限于缺乏细粒度的失真类型标注。为此,我们提出CAMP-VQA,一种基于大视觉语言模型语义理解能力的新框架。该方法引入质量感知提示机制,将视频元数据(如分辨率、帧率、码率)与从帧间变化中提取的关键片段结合,引导BLIP-2生成细粒度质量描述。设计统一架构,建模语义对齐、时间特性与空间特性三个维度的多模态特征,融合后回归得到视频质量分数。在多种UGC数据集上的实验表明,本模型显著优于现有方法,在无需人工细粒度标注的情况下实现更高精度,平均排名与线性相关性(SRCC: 0.928, PLCC: 0.938)均达最优。代码与训练模型已开源,并提供演示界面。

原文摘要 · Abstract (English)

The prevalence of user-generated content (UGC) on platforms such as YouTube and TikTok has rendered no-reference (NR) perceptual video quality assessment (VQA) vital for optimizing video delivery. Nonetheless, the characteristics of non-professional acquisition and the subsequent transcoding of UGC video on sharing platforms present significant challenges for NR-VQA. Although NR-VQA models attempt to infer mean opinion scores (MOS), their modeling of subjective scores for compressed content remains limited due to the absence of fine-grained perceptual annotations of artifact types. To address these challenges, we propose CAMP-VQA, a novel NR-VQA framework that exploits the semantic understanding capabilities of large vision-language models. Our approach introduces a quality-aware prompting mechanism that integrates video metadata (e.g., resolution, frame rate, bitrate) with key fragments extracted from inter-frame variations to guide the BLIP-2 pretraining approach in generating fine-grained quality captions. A unified architecture has been designed to model perceptual quality across three dimensions: semantic alignment, temporal characteristics, and spatial characteristics. These multimodal features are extracted and fused, then regressed to video quality scores. Extensive experiments on a wide variety of UGC datasets demonstrate that our model consistently outperforms existing NR-VQA methods, achieving improved accuracy without the need for costly manual fine-grained annotations. Our method achieves the best performance in terms of average rank and linear correlation (SRCC: 0.928, PLCC: 0.938) compared to state-of-the-art methods. The source code and trained models, along with a user-friendly demo, are available at: https://github.com/xinyiW915/CAMP-VQA.

视频质量评估多模态大模型无参考

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。