让多个角色按指定动作自然互动,实现文本生成视频的个性化定制。
VideoMage: Multi-Subject and Motion Customization of Text-to-Video Diffusion Models
- 用图像和视频分别提取主体身份与动作特征,解耦外观与运动
- 支持多主体在指定动作下自然交互,生成视频一致性高
- 适合需要多角色动态协作的视频创作场景
个性化文本到视频生成旨在生成融合用户指定主体身份或动作模式的高质量视频。然而,现有方法主要聚焦于单一概念的个性化,如主体身份或动作模式,难以有效处理多个主体同时具备特定动作模式的需求。为此,我们提出统一框架 VideoMage,实现对多个主体及其交互动作的视频定制。VideoMage 采用主体与动作 LoRA,从用户提供的图像和视频中捕捉个性化内容,并通过无外观依赖的动作学习方法,将动作模式与视觉外观解耦。此外,我们设计了时空组合机制,引导主体在目标动作模式下的自然交互。大量实验表明,VideoMage 在生成连贯、可控且主体身份一致的视频方面优于现有方法。
原文摘要 · Abstract (English)
Customized text-to-video generation aims to produce high-quality videos that incorporate user-specified subject identities or motion patterns. However, existing methods mainly focus on personalizing a single concept, either subject identity or motion pattern, limiting their effectiveness for multiple subjects with the desired motion patterns. To tackle this challenge, we propose a unified framework VideoMage for video customization over both multiple subjects and their interactive motions. VideoMage employs subject and motion LoRAs to capture personalized content from user-provided images and videos, along with an appearance-agnostic motion learning approach to disentangle motion patterns from visual appearance. Furthermore, we develop a spatial-temporal composition scheme to guide interactions among subjects within the desired motion patterns. Extensive experiments demonstrate that VideoMage outperforms existing methods, generating coherent, user-controlled videos with consistent subject identities and interactions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。