为文本生成视频质量评估打造新基准与评测方法。
T2VEval: Benchmark Dataset and Objective Evaluation Method for T2V-generated Videos
- 构建多维度数据集T2VEval-Bench,含148个提示和1783个生成视频。
- 提出T2VEval多分支融合模型,三维度评估并实现最优性能。
- 适合视频生成模型优化与质量评测的研究者使用。
近期以Runway Gen-3、Pika、Sora和Kling为代表的文本到视频(T2V)技术迅速发展,推动了该技术的广泛应用与普及。然而,由于存在如动作不自然、违背人类认知等复杂失真现象,对T2V输出的质量评估仍面临挑战。为此,我们构建了T2VEval-Bench——一个用于T2V质量评估的多维基准数据集,包含148个文本提示和1,783个由13个T2V模型生成的视频。为全面评估,我们在主观实验中从整体印象、文视频一致性、真实感和技术质量四个维度评分。基于此数据集,我们提出了T2VEval,一种多分支融合的T2V质量评估方法。T2VEval通过三个分支分别评估文视频一致性、真实感和技术质量,利用注意力融合模块整合各分支特征,并借助大语言模型预测最终得分。此外,采用分而治之的训练策略,使各分支在专注学习的同时保持协同。实验表明,T2VEval在多项指标上达到当前最佳表现。
原文摘要 · Abstract (English)
Recent advances in text-to-video (T2V) technology, as demonstrated by models such as Runway Gen-3, Pika, Sora, and Kling, have significantly broadened the applicability and popularity of the technology. This progress has created a growing demand for accurate quality assessment metrics to evaluate the perceptual quality of T2V-generated videos and optimize video generation models. However, assessing the quality of text-to-video outputs remain challenging due to the presence of highly complex distortions, such as unnatural actions and phenomena that defy human cognition. To address these challenges, we constructed T2VEval-Bench, a multi-dimensional benchmark dataset for text-to-video quality evaluation, which contains 148 textual prompts and 1,783 videos generated by 13 T2V models. To ensure a comprehensive evaluation, we scored each video on four dimensions in the subjective experiment, which are overall impression, text-video consistency, realness, and technical quality. Based on T2VEval-Bench, we developed T2VEval, a multi-branch fusion scheme for T2V quality evaluation. T2VEval assesses videos across three branches: text-video consistency, realness, and technical quality. Using an attention-based fusion module, T2VEval effectively integrates features from each branch and predicts scores with the aid of a large language model. Additionally, we implemented a divide-and-conquer training strategy, enabling each branch to learn targeted knowledge while maintaining synergy with the others. Experimental results demonstrate that T2VEval achieves state-of-the-art performance across multiple metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。