arXiv:2412.18060cs.CV2024-12中稿 · ICASSP 2025被引 8

用多模态大模型提升短视频质量评估,效果优于传统方法

An Ensemble Approach to Short-form Video Quality Assessment Using Multimodal LLM

  • 结合多模态大模型与现有质量评估模型,构建自适应集成方法
  • 在多个数据集上实现更优的泛化性能,尤其在复杂编辑视频中表现突出
  • 适合关注短视频质量评估、多模态模型应用的研究者与工程师

短时视频内容多样、剪辑风格丰富且常含视觉伪影,给基于学习的盲视频质量评估(BVQA)模型带来巨大挑战。多模态大语言模型(MLLM)凭借出色的泛化能力,展现出解决该问题的潜力。本文研究如何有效利用预训练的MLLM进行短时视频质量评估,重点分析了预处理与响应变异的影响,并探索将MLLM与现有BVQA模型融合的策略。首先,我们考察了帧预处理和采样方式对MLLM性能的影响;随后提出一种轻量级学习型集成方法,自适应融合MLLM与先进BVQA模型的预测结果。实验表明,所提集成方法具有优异的泛化能力。此外,内容感知的集成权重分析揭示:部分视频特征未被现有BVQA模型充分表达,提示了未来改进方向。

原文摘要 · Abstract (English)

The rise of short-form videos, characterized by diverse content, editing styles, and artifacts, poses substantial challenges for learning-based blind video quality assessment (BVQA) models. Multimodal large language models (MLLMs), renowned for their superior generalization capabilities, present a promising solution. This paper focuses on effectively leveraging a pretrained MLLM for short-form video quality assessment, regarding the impacts of pre-processing and response variability, and insights on combining the MLLM with BVQA models. We first investigated how frame pre-processing and sampling techniques influence the MLLM's performance. Then, we introduced a lightweight learning-based ensemble method that adaptively integrates predictions from the MLLM and state-of-the-art BVQA models. Our results demonstrated superior generalization performance with the proposed ensemble approach. Furthermore, the analysis of content-aware ensemble weights highlighted that some video characteristics are not fully represented by existing BVQA models, revealing potential directions to improve BVQA models further.

视频质量评估多模态大模型集成学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。