构建首个图文生成视频的时空伪影标注数据集
GeneVA: A Dataset of Human Annotations for Generative Text to Video Artifacts
- 基于自然文本提示生成视频,标注时空伪影
- 涵盖不可实现物理与时间不一致等典型问题
- 适用于模型评估与生成质量优化
近年来,概率生成模型已从静态图像合成扩展到文本驱动的视频生成。然而,其生成过程固有的随机性可能导致不可预测的伪影,如违背物理规律和时间不一致等问题。解决这些挑战需要系统性的基准测试,但现有数据集主要聚焦于生成图像,因视频特有的时空复杂性而受限。为此,我们提出GeneVA,一个大规模、富含人工标注的伪影数据集,专注于从自然文本提示生成的视频中的时空伪影。我们期望GeneVA能推动模型性能评估与生成视频质量提升等关键应用。
原文摘要 · Abstract (English)
Recent advances in probabilistic generative models have extended capabilities from static image synthesis to text-driven video generation. However, the inherent randomness of their generation process can lead to unpredictable artifacts, such as impossible physics and temporal inconsistency. Progress in addressing these challenges requires systematic benchmarks, yet existing datasets primarily focus on generative images due to the unique spatio-temporal complexities of videos. To bridge this gap, we introduce GeneVA, a large-scale artifact dataset with rich human annotations that focuses on spatio-temporal artifacts in videos generated from natural text prompts. We hope GeneVA can enable and assist critical applications, such as benchmarking model performance and improving generative video quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。