无需微调即可实现多概念视频个性化定制,解决身份混淆问题
ConceptMaster: Multi-Concept Video Customization on Diffusion Transformer Models Without Test-Time Tuning
- 分离学习多概念嵌入并独立注入扩散模型
- 在六种组合场景下均显著提升概念保真度与身份区分度
- 适合需要高效定制视频内容的研究者与开发者
文本到视频生成在扩散模型推动下取得显著进展,但多概念视频定制(MCVC)仍是重大挑战。本文指出两大关键问题:一是身份耦合问题,现有定制方法在处理多个概念时会混淆身份特征;二是高质量视频-实体配对数据稀缺,制约了模型对多样化概念的准确表达与解耦能力。为此,我们提出ConceptMaster框架,通过分离学习多概念嵌入,并以独立方式注入扩散模型,有效保障多身份视频的生成质量,即使面对视觉相似的概念也能保持高保真。为缓解数据稀缺问题,我们构建了一套数据生成流程,覆盖多种场景的高质量多概念视频-实体配对数据。同时设计了一个多概念视频评估集,从概念保真度、身份解耦能力与视频生成质量三个维度,在六种不同概念组合场景下全面验证方法。大量实验表明,ConceptMaster显著优于先前方法,展现出在视频扩散模型中生成个性化、语义精准内容的巨大潜力。
原文摘要 · Abstract (English)
Text-to-video generation has made remarkable advancements through diffusion models. However, Multi-Concept Video Customization (MCVC) remains a significant challenge. We identify two key challenges for this task: 1) the identity decoupling issue, where directly adopting existing customization methods inevitably mix identity attributes when handling multiple concepts simultaneously, and 2) the scarcity of high-quality video-entity pairs, which is crucial for training a model that can well represent and decouple various customized concepts in video generation. To address these challenges, we introduce ConceptMaster, a novel framework that effectively addresses the identity decoupling issues while maintaining concept fidelity in video customization. Specifically, we propose to learn decoupled multi-concept embeddings and inject them into diffusion models in a standalone manner, which effectively guarantees the quality of customized videos with multiple identities, even for highly similar visual concepts. To overcome the scarcity of high-quality MCVC data, we establish a data construction pipeline, which enables collection of high-quality multi-concept video-entity data pairs across diverse scenarios. A multi-concept video evaluation set is further devised to comprehensively validate our method from three dimensions, including concept fidelity, identity decoupling ability, and video generation quality, across six different concept composition scenarios. Extensive experiments demonstrate that ConceptMaster significantly outperforms previous methods for video customization tasks, showing great potential to generate personalized and semantically accurate content for video diffusion models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。