构建6万+视频评估数据集,用大模型实现文本-视频生成与理解的多维度自动评测。
LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation
- 基于大模型设计多维度评估指标LOVE,覆盖视觉偏好、图文对齐等
- 在6万+视频上构建超大规模人工标注数据集,含12万评分和6万问答对
- 可同时评估生成与理解能力,适合研究者与开发者验证AIGV模型
近年来,大型多模态模型(LMMs)推动了文本到视频(T2V)生成和视频到文本(V2T)理解任务的显著进展。然而,当前人工智能生成视频(AIGVs)在感知质量与图文一致性方面仍存在局限。因此,亟需一个可靠且可扩展的自动化评估方法,这高度依赖于高质量的人工标注规模。为此,我们提出AIGVE-60K,一个全面的AI生成视频评估数据集与基准,包含:(i) 覆盖20个细粒度任务维度的3,050个详细提示;(ii) 最大规模的人工标注,包括120,000个平均意见得分(MOS)和60,000个问答对,标注于58,500段由30个T2V模型生成的视频;(iii) 双向基准测试,用于评估T2V生成与V2T理解能力。基于AIGVE-60K,我们提出LOVE——一种基于大模型的多维度评估指标,涵盖感知偏好、图文对应关系及任务特定准确率,支持实例级与模型级评估。大量实验表明,LOVE不仅在AIGVE-60K上达到领先性能,还有效泛化至多种其他AIGV评估基准。这些结果凸显了AIGVE-60K的重要性。数据库与代码已匿名公开于https://github.com/IntMeGroup/LOVE。
原文摘要 · Abstract (English)
Recent advancements in large multimodal models (LMMs) have driven substantial progress in both text-to-video (T2V) generation and video-to-text (V2T) interpretation tasks. However, current AI-generated videos (AIGVs) still exhibit limitations in terms of perceptual quality and text-video alignment. Therefore, a reliable and scalable automatic model for AIGV evaluation is desirable, which heavily relies on the scale and quality of human annotations. To this end, we present AIGVE-60K, a comprehensive dataset and benchmark for AI-Generated Video Evaluation, which features (i) comprehensive tasks, encompassing 3,050 extensive prompts across 20 fine-grained task dimensions, (ii) the largest human annotations, including 120K mean-opinion scores (MOSs) and 60K question-answering (QA) pairs annotated on 58,500 videos generated from 30 T2V models, and (iii) bidirectional benchmarking and evaluating for both T2V generation and V2T interpretation capabilities. Based on AIGVE-60K, we propose LOVE, a LMM-based metric for AIGV Evaluation from multiple dimensions including perceptual preference, text-video correspondence, and task-specific accuracy in terms of both instance level and model level. Comprehensive experiments demonstrate that LOVE not only achieves state-of-the-art performance on the AIGVE-60K dataset, but also generalizes effectively to a wide range of other AIGV evaluation benchmarks. These findings highlight the significance of the AIGVE-60K dataset. Database and codes are anonymously available at https://github.com/IntMeGroup/LOVE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。