arXiv:2506.13691cs.CV2025-06NeurIPS被引 55

构建首个支持4K/8K的高质量图文视频数据集,助力超高清视频生成研究。

UltraVideo: High-Quality UHD Video Dataset with Comprehensive Captions

  • 四阶段自动化流程筛选高质量视频并生成9条结构化描述
  • 数据集含100+主题、22.4%为8K分辨率,每视频平均824词摘要
  • 适配高分辨率视频生成模型训练,推动影视级视频合成发展

视频数据集的质量(图像质量、分辨率、细粒度描述)显著影响视频生成模型的表现。随着视频应用对高质量生成需求提升,如电影级超高清(UHD)视频与4K短视频内容生成,现有公开数据集已无法满足研究与应用需求。本文首次提出一个开源的高质量文本到视频数据集UltraVideo,支持UHD-4K(其中22.4%为8K),涵盖100多种主题,每个视频配备9条结构化描述及1条总结性描述(平均824词)。我们设计了四阶段高度自动化的数据清洗流程:(i)多样化高质量视频片段收集;(ii)基于统计特征过滤;(iii)模型驱动的数据净化;(iv)生成全面且结构化的标题。此外,我们扩展了Wan模型为UltraWan-1K/-4K,可原生生成高质量1K/4K视频,且文本控制更一致,验证了数据清洗的有效性。UltraVideo数据集与UltraWan模型已开放获取。

原文摘要 · Abstract (English)

The quality of the video dataset (image quality, resolution, and fine-grained caption) greatly influences the performance of the video generation model. The growing demand for video applications sets higher requirements for high-quality video generation models. For example, the generation of movie-level Ultra-High Definition (UHD) videos and the creation of 4K short video content. However, the existing public datasets cannot support related research and applications. In this paper, we first propose a high-quality open-sourced UHD-4K (22.4\% of which are 8K) text-to-video dataset named UltraVideo, which contains a wide range of topics (more than 100 kinds), and each video has 9 structured captions with one summarized caption (average of 824 words). Specifically, we carefully design a highly automated curation process with four stages to obtain the final high-quality dataset: \textit{i)} collection of diverse and high-quality video clips. \textit{ii)} statistical data filtering. \textit{iii)} model-based data purification. \textit{iv)} generation of comprehensive, structured captions. In addition, we expand Wan to UltraWan-1K/-4K, which can natively generate high-quality 1K/4K videos with more consistent text controllability, demonstrating the effectiveness of our data curation.We believe that this work can make a significant contribution to future research on UHD video generation. UltraVideo dataset and UltraWan models are available at https://xzc-zju.github.io/projects/UltraVideo.

视频生成超高清数据集图文对

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。