arXiv:2506.15838cs.CV2025-06NeurIPS被引 27

让AI一次生成多镜头一致人像视频,支持个性化定制。

EchoShot: Multi-Shot Portrait Video Generation

  • 用新位置编码建模多镜头差异,直接训练多帧视频
  • 在高保真人像数据集上实现身份与属性精准一致
  • 适合影视创作、虚拟形象等需要多角度连贯视频的场景

视频扩散模型显著提升了艺术创作效率,具备高质量人像视频生成能力。然而,现有方法主要局限于单镜头生成,而实际应用亟需保持身份一致性且可控内容的多镜头视频。本文提出EchoShot,一种基于基础视频扩散模型的原生可扩展多镜头人像生成框架。首先,在视频扩散变压器架构中引入镜头感知的位置嵌入机制,以建模镜头间变化,并建立多镜头视觉内容与文本描述之间的精细对应关系。该设计简单有效,无需额外计算开销即可直接在多镜头视频数据上训练。为支持多镜头场景下的模型训练,构建了PortraitGala——一个大规模、高保真的人像中心视频数据集,包含跨镜头身份一致性与细粒度标注(如面部特征、服饰、动态动作)。进一步拓展EchoShot,实现基于参考图像的个性化多镜头生成及无限镜头数量的长视频合成。大量评估表明,EchoShot在多镜头人像视频生成中表现出更优的身份一致性与属性级可控性,展现出作为通用多镜头视频建模基础范式的潜力。

原文摘要 · Abstract (English)

Video diffusion models substantially boost the productivity of artistic workflows with high-quality portrait video generative capacity. However, prevailing pipelines are primarily constrained to single-shot creation, while real-world applications urge for multiple shots with identity consistency and flexible content controllability. In this work, we propose EchoShot, a native and scalable multi-shot framework for portrait customization built upon a foundation video diffusion model. To start with, we propose shot-aware position embedding mechanisms within video diffusion transformer architecture to model inter-shot variations and establish intricate correspondence between multi-shot visual content and their textual descriptions. This simple yet effective design enables direct training on multi-shot video data without introducing additional computational overhead. To facilitate model training within multi-shot scenario, we construct PortraitGala, a large-scale and high-fidelity human-centric video dataset featuring cross-shot identity consistency and fine-grained captions such as facial attributes, outfits, and dynamic motions. To further enhance applicability, we extend EchoShot to perform reference image-based personalized multi-shot generation and long video synthesis with infinite shot counts. Extensive evaluations demonstrate that EchoShot achieves superior identity consistency as well as attribute-level controllability in multi-shot portrait video generation. Notably, the proposed framework demonstrates potential as a foundational paradigm for general multi-shot video modeling.

视频生成多镜头身份一致扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。