评测大模型对视频美学的感知能力,发现当前模型表现基础且不精准。
VideoAesBench: Benchmarking the Video Aesthetics Perception Capabilities of Large Multimodal Models
- 构建多源视频数据集,涵盖用户生成、AI生成等五类视频。
- 设计多类型题目,包含开放描述题,覆盖视觉形式、风格与情感三维度。
- 测试23个主流模型,揭示其在视频美学理解上仍处初级阶段。
大型多模态模型(LMMs)在诸多视觉感知任务中表现出色,但其对视频美学质量的评估能力——人类基本认知能力之一——仍未得到充分研究。为此,我们提出VideoAesBench,一个全面评估LMMs视频美学理解能力的基准。该基准具有三大特征:(1) 多样化内容,包含来自用户生成(UGC)、AI生成(AIGC)、压缩视频、机器人生成(RGC)及游戏视频等来源的1,804段视频;(2) 多样化题型,包括传统单选、多选、判断题,以及新颖的开放式美学描述题;(3) 全面评估维度,涵盖视觉形式(5个方面)、视觉风格(4个方面)和视觉感染力(3个方面)。基于此,我们对23个开源与商业LMM进行了评测。结果表明,当前模型仅具备基础的视频美学感知能力,性能不完整且不精确。我们希望VideoAesBench能成为强效测评平台,推动可解释性视频美学评估的发展。数据将发布于https://github.com/michaelliyunhao/VideoAesBench。
原文摘要 · Abstract (English)
Large multimodal models (LMMs) have demonstrated outstanding capabilities in various visual perception tasks, which has in turn made the evaluation of LMMs significant. However, the capability of video aesthetic quality assessment, which is a fundamental ability for human, remains underexplored for LMMs. To address this, we introduce VideoAesBench, a comprehensive benchmark for evaluating LMMs' understanding of video aesthetic quality. VideoAesBench has several significant characteristics: (1) Diverse content including 1,804 videos from multiple video sources including user-generated (UGC), AI-generated (AIGC), compressed, robotic-generated (RGC), and game videos. (2) Multiple question formats containing traditional single-choice questions, multi-choice questions, True or False questions, and a novel open-ended questions for video aesthetics description. (3) Holistic video aesthetics dimensions including visual form related questions from 5 aspects, visual style related questions from 4 aspects, and visual affectiveness questions from 3 aspects. Based on VideoAesBench, we benchmark 23 open-source and commercial large multimodal models. Our findings show that current LMMs only contain basic video aesthetics perception ability, their performance remains incomplete and imprecise. We hope our VideoAesBench can be served as a strong testbed and offer insights for explainable video aesthetics assessment. The data will be released on https://github.com/michaelliyunhao/VideoAesBench
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。