打造百万级定制视频生成数据集与模型,支持千类真实场景应用。
A Comprehensive Ecosystem for Open-Domain Customized Video Generation

- 构建百万级<身份, 文本, 视频>三元组数据集PexelsCustom-1M
- 仅用8%额外参数实现高效定制化视频生成,性能超越现有方法
- 推出千类规模新基准OpenCustom,适配真实世界多场景需求
视频生成技术虽取得显著进展,但开放域定制化视频生成受限于缺乏大规模、带标注的多样化身份属性数据集。为此,我们提出首个公开可用的百万级数据集PexelsCustom-1M,包含超过100万条经筛选的<身份, 文本, 视频>三元组,覆盖8000+类别。基于此,我们设计了CustoMDiT框架,通过仅增加8%可学习参数,将预训练多模态Diffusion Transformer高效适配为定制化视频生成器。该方法在性能上超越现有最先进水平。传统基准如DreamBooth仅覆盖100类,难以满足实际应用需求;为此,我们通过融合ImageNet与MS-COCO构建了包含1000+类别的新基准OpenCustom。大量实验验证了数据集与模型的优势。我们将开源整个生态系统——包括数据集、生成流程、基准和代码实现,以推动后续研究。
原文摘要 · Abstract (English)
Recent progress in video generation has shown impressive visual synthesis capabilities. However, open-domain customized video generation remains limited by the lack of large-scale, annotated datasets capturing diverse identity-specific attributes. To address this, we introduce PexelsCustom-1M, the first publicly available million-scale dataset for identity-preserving video generation, containing one million curated <identity, text, video> triplets across 8,000+ categories. Leveraging this, we propose CustoMDiT, a parameter-efficient framework that adapts a pretrained multimodal Diffusion Transformer into a customized video generator with only 8% additional learnable parameters. Our method surpasses prior state-of-the-art. However, benchmarks such as DreamBooth cover only 100 classes, which is insufficient for real-world applications. To overcome this, we construct OpenCustom, a new benchmark with 1,000+ categories, created via cross-dataset knowledge fusion from ImageNet and MS-COCO. Extensive experiments confirm the advantages of both our dataset and model. We will open-source the entire ecosystem--including dataset, pipeline, benchmark, and implementations--to support further research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。