arXiv:2506.18851cs.CV2025-06被引 18

构建首个通用跨对一致性视频生成数据集,解决文本指令跟随难题

Phantom-Data : Towards a General Subject-Consistent Video Generation Dataset

  • 通过三阶段流程构建百万级跨场景身份一致视频对
  • 训练后模型在身份一致性上媲美传统方法,且更符合文本指令
  • 适合需要精准控制人物/物体身份的视频生成研究者

近年来,主体到视频生成取得了显著进展,但现有模型仍难以忠实遵循文本指令,这一问题被称为“复制粘贴”问题,源于广泛采用的成对训练范式。该方法通过从同一场景中采样参考图像,将主体身份与背景和上下文属性固有地纠缠在一起。为解决此问题,我们提出了首个通用跨对主体一致性视频生成数据集 Phantom-Data,包含约一百万个体感一致的数据对,覆盖多种类别。该数据集通过三阶段流程构建:(1) 通用且输入对齐的主体检测模块;(2) 从超过5300万视频和30亿图像中进行大规模跨上下文主体检索;(3) 基于先验的视觉一致性验证,以确保在不同上下文中保持身份一致。综合实验表明,使用 Phantom-Data 训练可显著提升提示对齐性和视觉质量,同时在身份一致性方面达到与成对基线相当的水平。

原文摘要 · Abstract (English)

Subject-to-video generation has witnessed substantial progress in recent years. However, existing models still face significant challenges in faithfully following textual instructions. This limitation, commonly known as the copy-paste problem, arises from the widely used in-pair training paradigm. This approach inherently entangles subject identity with background and contextual attributes by sampling reference images from the same scene as the target video. To address this issue, we introduce \textbf{Phantom-Data, the first general-purpose cross-pair subject-to-video consistency dataset}, containing approximately one million identity-consistent pairs across diverse categories. Our dataset is constructed via a three-stage pipeline: (1) a general and input-aligned subject detection module, (2) large-scale cross-context subject retrieval from more than 53 million videos and 3 billion images, and (3) prior-guided identity verification to ensure visual consistency under contextual variation. Comprehensive experiments show that training with Phantom-Data significantly improves prompt alignment and visual quality while preserving identity consistency on par with in-pair baselines.

视频生成身份一致性数据集文本对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。