构建视频美学与生成质量联合评估新基准,支持模型自动优化。
VGA-BenchV2: An Expanded Unified Benchmark and Multi-Model Framework for Evaluating Video Aesthetics and Generation Quality

- 基于52个细粒度维度构建统一评测框架
- 新增3.6万条人工标注,涵盖美学与生成质量
- 支持从评估到优化的闭环训练,提升视频美感
我们提出VGA-BenchV2,一个扩展的人类对齐基准与优化框架,用于联合评估和提升视频生成质量与美学价值。基于原有VGA-Bench,VGA-BenchV2保留了美学与生成两大主维度及52个子维度。依据该分类体系,我们收集了1,016个多样化提示词,并生成超过60,000段由12个主流视频生成模型产出的视频。更重要的是,VGA-BenchV2大幅扩充了人工标注数据,新增36,000条任务级标签,包括16,200条美学评分、13,200条美学标签、6,600条生成质量评估,分别较VGA-Bench提升13.46倍、11.15倍和1.55倍。利用这一扩大化的标注语料,我们设计了混合评估架构:VAQA-Net用于连续美学打分,两个基于Qwen的视觉语言模型VTag-Net和VGQA-Net分别用于美学标签识别与生成质量判断。大量实验表明其评估结果与人类判断高度一致。除评估外,VGA-BenchV2进一步引入评估-优化闭环流程,将学习到的美学评估器作为强化学习中的奖励模型,用于生成器微调。该流程实现从基准构建、人工标注到自动化评估与模型优化的全链条闭环,使视频生成模型不仅能提升真实感,还能增强美学品质与人类偏好契合度。资源已公开于https://huggingface.co/datasets/BestiVictoryLab/VGA-Bench。
原文摘要 · Abstract (English)
We introduce VGA-BenchV2, an extended human-aligned benchmark and optimization framework for jointly evaluating and improving video generation quality and aesthetic value. Built upon VGA-Bench, VGA-BenchV2 preserves the original fine-grained taxonomy with two primary dimensions-Aesthetic and Generation-and 52 sub-dimensions. Guided by this taxonomy, we curate 1,016 diverse prompts and collect over 60,000 videos generated by 12 mainstream video generation models. More importantly, VGA-BenchV2 substantially expands human-labeled supervision by adding 36,000 task-level annotations, including 16,200 for aesthetic quality, 13,200 for aesthetic tagging, and 6,600 for generation quality, corresponding to 13.46x, 11.15x, and 1.55x scale-ups over VGA-Bench, respectively. Leveraging this enlarged annotation corpus, we develop a hybrid evaluator architecture consisting of VAQA-Net for continuous aesthetic scoring and two Qwen-based Large Vision-Language Model evaluators, VTag-Net and VGQA-Net, for aesthetic tagging and generation quality assessment. Extensive experiments demonstrate strong alignment with human judgments across diverse generation models. Beyond evaluation, VGA-BenchV2 further introduces an evaluation-to-optimization pipeline, where the learned aesthetic evaluator serves as a reward model for reinforcement learning-based generator fine-tuning. This closes the loop from benchmark construction and human supervision to automated evaluation and model optimization, enabling video generators to improve not only in realism but also in aesthetic quality and human preference alignment. Resources are available at https://huggingface.co/datasets/BestiVictoryLab/VGA-Bench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。