首个聚焦用户需求的百万级文生视频数据集,解决生成内容与实际需求不匹配问题。
VideoUFO: A Million-Scale User-Focused Dataset for Text-to-Video Generation

- 从真实用户提问中提炼1291个主题,精准匹配用户关注点。
- 构建超百万视频片段数据集,覆盖广泛且与现有数据重合率仅0.29%。
- 适合训练更贴近真实需求的文生视频模型,尤其提升冷门主题表现。
文生视频模型将文本提示转化为动态视觉内容,在影视制作、游戏和教育等领域有广泛应用。然而其实际表现常低于用户期望,主要因缺乏相关主题的训练数据。本文提出VideoUFO,首个专为用户关注点设计的视频数据集。该数据集具有:(1)与现有视频数据集仅有0.29%重合;(2)所有视频均通过YouTube官方API、在创作共用许可下获取。数据集包含超过109万视频片段,每段配有简短与详细双标签描述。我们首先从百万级真实文生视频提示数据集VidProM中聚类出1291个用户关注主题,再基于这些主题检索YouTube视频,切分为片段,并生成双层描述。经主题验证后保留约109万片段。实验表明:(1)当前16个文生视频模型在不同用户主题上表现不一致;(2)仅用VideoUFO训练的简单模型,在表现最差的主题上超越其他模型。数据与代码已开源,许可协议为CC BY 4.0。
原文摘要 · Abstract (English)
Text-to-video generative models convert textual prompts into dynamic visual content, offering wide-ranging applications in film production, gaming, and education. However, their real-world performance often falls short of user expectations. One key reason is that these models have not been trained on videos related to some topics users want to create. In this paper, we propose VideoUFO, the first Video dataset specifically curated to align with Users' FOcus in real-world scenarios. Beyond this, our VideoUFO also features: (1) minimal (0.29%) overlap with existing video datasets, and (2) videos searched exclusively via YouTube's official API under the Creative Commons license. These two attributes provide future researchers with greater freedom to broaden their training sources. The VideoUFO comprises over 1.09 million video clips, each paired with both a brief and a detailed caption (description). Specifically, through clustering, we first identify 1,291 user-focused topics from the million-scale real text-to-video prompt dataset, VidProM. Then, we use these topics to retrieve videos from YouTube, split the retrieved videos into clips, and generate both brief and detailed captions for each clip. After verifying the clips with specified topics, we are left with about 1.09 million video clips. Our experiments reveal that (1) current 16 text-to-video models do not achieve consistent performance across all user-focused topics; and (2) a simple model trained on VideoUFO outperforms others on worst-performing topics. The dataset and code are publicly available at https://huggingface.co/datasets/WenhaoWang/VideoUFO and https://github.com/WangWenhao0716/BenchUFO under the CC BY 4.0 License.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。