arXiv:2604.10127cs.CVcs.AI2026-04被引 2

构建首个视频美学与生成质量联合评估基准,填补艺术性评价空白。

VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation

论文配图:VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation
图 1 · 摘自论文原文
  • 基于三层细粒度分类体系,系统化定义视频美学与生成质量维度。
  • 生成超6万段视频数据,覆盖12个模型、1016个多样化提示词。
  • 推出三类神经评估器,实现自动美学评分、标签识别与质量判断。

AIGC视频生成技术的快速发展凸显了超越传统生成质量指标、综合评估美学吸引力的迫切需求。现有基准多聚焦技术保真度,忽视感知与艺术性等整体评估。为此,我们提出VGA-Bench,一个统一的视频生成质量与美学质量联合评估基准。该基准基于三层次分类体系:美学质量、美学标签、生成质量,每项分解为多个细粒度子维度,实现系统性评估。基于此框架,设计1,016个多样化提示词,使用12个视频生成模型生成超过6万段视频,覆盖丰富内容、风格与伪影。为实现可扩展的自动化评估,通过人工标注子集并开发三类专用多任务神经评估器:VAQA-Net用于美学质量预测,VTag-Net用于自动美学标签,VGQA-Net用于生成与基础质量属性评估。大量实验表明,模型与人类判断高度一致,兼具准确与高效。我们公开发布VGA-Bench,以推动AIGC评估研究,应用于内容审核、模型调试与生成模型优化。

原文摘要 · Abstract (English)

The rapid advancement of AIGC-based video generation has underscored the critical need for comprehensive evaluation frameworks that go beyond traditional generation quality metrics to encompass aesthetic appeal. However, existing benchmarks remain largely focused on technical fidelity, leaving a significant gap in holistic assessment-particularly with respect to perceptual and artistic qualities. To address this limitation, we introduce VGA-Bench, a unified benchmark for joint evaluation of video generation quality and aesthetic quality. VGA-Bench is built upon a principled three-tier taxonomy: Aesthetic Quality, Aesthetic Tagging, and Generation Quality, each decomposed into multiple fine-grained sub-dimensions to enable systematic assessment. Guided by this taxonomy, we design 1,016 diverse prompts and generate a large-scale dataset of over 60,000 videos using 12 video generation models, ensuring broad coverage across content, style, and artifacts. To enable scalable and automated evaluation, we annotate a subset of the dataset via human labeling and develop three dedicated multi-task neural assessors: VAQA-Net for aesthetic quality prediction, VTag-Net for automatic aesthetic tagging, and VGQA-Net for generation and basic quality attributes. Extensive experiments demonstrate that our models achieve reliable alignment with human judgments, offering both accuracy and efficiency. We release VGA-Bench as a public benchmark to foster research in AIGC evaluation, with applications in content moderation, model debugging, and generative model optimization.

视频评估美学质量AIGC

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。