arXiv:2409.20063cs.CV2024-09CVPR被引 44

评测大模型对视频质量的理解能力,发现其表现远低于人类。

Q-Bench-Video: Benchmarking the Video Quality Understanding of LMMs

  • 构建涵盖真实、AI生成和电脑动画的视频质量评测集
  • 包含2378个问答对,测试12个开源与5个专有模型
  • 首次引入AIGC质量失真维度,推动视频生成研究

随着大型多模态模型(LMMs)在视频理解领域的兴起,现有研究多聚焦于通用视频理解能力,忽视了对视频质量理解的系统性探索。为此,本文提出Q-Bench-Video,一个专门评估LMMs视频质量理解能力的新基准。该基准涵盖自然场景、AI生成内容(AIGC)和计算机图形(CG)三类视频源,采用多选题、是非题、如何题及开放题形式,并新增视频对质量对比题以提升全面性。评估维度扩展至技术、美学、时间与AIGC失真四个方面,以应对日益增长的视频生成需求。共收集2,378个问答对,在12个开源与5个专有LMMs上进行测试。结果表明,尽管LMMs具备一定的视频质量理解基础,但整体表现仍不完整且不精确,与人类表现存在显著差距。Q-Bench-Video旨在激发社区关注,推动相关研究,挖掘LMMs在视频质量理解方面的潜力。

原文摘要 · Abstract (English)

With the rising interest in research on Large Multi-modal Models (LMMs) for video understanding, many studies have emphasized general video comprehension capabilities, neglecting the systematic exploration into video quality understanding. To address this oversight, we introduce Q-Bench-Video in this paper, a new benchmark specifically designed to evaluate LMMs' proficiency in discerning video quality. a) To ensure video source diversity, Q-Bench-Video encompasses videos from natural scenes, AI-generated Content (AIGC), and Computer Graphics (CG). b) Building on the traditional multiple-choice questions format with the Yes-or-No and What-How categories, we include Open-ended questions to better evaluate complex scenarios. Additionally, we incorporate the video pair quality comparison question to enhance comprehensiveness. c) Beyond the traditional Technical, Aesthetic, and Temporal distortions, we have expanded our evaluation aspects to include the dimension of AIGC distortions, which addresses the increasing demand for video generation. Finally, we collect a total of 2,378 question-answer pairs and test them on 12 open-source & 5 proprietary LMMs. Our findings indicate that while LMMs have a foundational understanding of video quality, their performance remains incomplete and imprecise, with a notable discrepancy compared to human performance. Through Q-Bench-Video, we seek to catalyze community interest, stimulate further research, and unlock the untapped potential of LMMs to close the gap in video quality understanding.

视频质量多模态AIGC评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。