arXiv:2505.16025cs.CVcs.MM2025-05中稿 · ICIP 2026被引 4

CP-LLM融合上下文与像素信息,精准评估视频质量并生成可解释描述。

Context and Pixel Aware Large Language Model for Video Quality Assessment

  • 双视觉编码器分别处理视频上下文与像素级失真
  • 在多个数据集上达到顶尖性能,对压缩伪影更敏感
  • 适合需要高质量评分与解释的视频分析场景

视频质量评估(VQA)是具有广泛应用价值的挑战性课题。传统手工设计和判别式学习模型主要关注像素级失真,缺乏上下文理解;而近期多模态大语言模型(MLLMs)对微小失真不敏感,且将质量评分与描述视为独立任务。为此,我们提出CP-LLM:一种上下文与像素感知的大语言模型。该模型采用双视觉编码器,分别以高层(视频上下文)和低层(像素失真)粒度独立分析感知质量,并通过语言解码器推理两者间关系。此设计使CP-LLM能同时输出鲁棒的质量评分与可解释的质量描述,对压缩伪影等微小失真更敏感。实验表明,CP-LLM在多个VQA基准上实现跨数据集最优性能,且对像素失真表现出更强鲁棒性。

原文摘要 · Abstract (English)

Video quality assessment (VQA) is a challenging research topic with broad applications. Traditional hand-crafted and discriminative learning-based VQA models mainly focus on pixel-level distortions and lack contextual understanding, while recent multimodal large language models (MLLMs) struggle with sensitivity to small distortions or handle quality scoring and description as separate tasks. To address these shortcomings, we introduce CP-LLM: a Context- and Pixel-aware Large Language Model. CP-LLM is a novel multimodal LLM architecture featuring dual vision encoders designed to independently analyze perceptual quality at both high-level (video context) and low-level (pixel distortion) granularity, along with a language decoder that subsequently reasons about the interplay between these aspects. This design enables CP-LLM to simultaneously produce robust quality scores and interpretable quality descriptions, with enhanced sensitivity to pixel distortions (e.g., compression artifacts). Experiment results demonstrate that CP-LLM achieves state-of-the-art cross-dataset performance on VQA benchmarks and superior robustness to pixel distortions.

视频质量评估多模态大模型感知建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。