用锚点提示实现多概念视频个性化,不调参也能保持身份清晰
Movie Weaver: Tuning-Free Multi-Concept Video Personalization with Anchored Prompts
- 用图像锚点作为唯一文本标记,精准关联参考图与生成内容
- 支持人脸、身体、动物等多概念组合,生成视频身份保留率更高
- 无需微调模型,适合快速生成多角色定制视频的场景
视频个性化通过参考图像生成定制化视频,近年来受到广泛关注。然而,现有方法多局限于单概念个性化,难以实现多概念融合;扩展至多概念时易出现身份混杂,导致生成角色融合多个来源特征。其根源在于缺乏将每个概念与特定参考图像关联的机制。本文提出锚点提示(anchored prompts),将图像锚点嵌入文本提示中作为唯一标识,引导生成过程准确引用对应参考图。同时引入概念嵌入以编码参考图像的顺序。所提出的 Movie Weaver 方法可无缝整合人脸、身体、动物等多类参考图像生成统一视频,实现灵活组合且无需模型微调。评估结果表明,该方法在身份保留和整体质量上均优于现有技术。
原文摘要 · Abstract (English)
Video personalization, which generates customized videos using reference images, has gained significant attention. However, prior methods typically focus on single-concept personalization, limiting broader applications that require multi-concept integration. Attempts to extend these models to multiple concepts often lead to identity blending, which results in composite characters with fused attributes from multiple sources. This challenge arises due to the lack of a mechanism to link each concept with its specific reference image. We address this with anchored prompts, which embed image anchors as unique tokens within text prompts, guiding accurate referencing during generation. Additionally, we introduce concept embeddings to encode the order of reference images. Our approach, Movie Weaver, seamlessly weaves multiple concepts-including face, body, and animal images-into one video, allowing flexible combinations in a single model. The evaluation shows that Movie Weaver outperforms existing methods for multi-concept video personalization in identity preservation and overall quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。