让视频中两个人物的关系更精准可控,支持复杂互动生成。
DreamRelation: Relation-Centric Video Customization
- 用关系解耦学习分离人物动作与互动关系,提升泛化能力。
- 引入时空关系对比损失,聚焦动态交互而非细节外观。
- 首次实现可解释的关系生成框架,适合影视创作与虚拟人开发。
关系型视频定制旨在生成体现用户指定两个主体间关系的个性化视频,对理解真实视觉内容至关重要。现有方法虽能个性化主体外观和动作,但在复杂关系定制上仍面临挑战,需精确建模关系并具备跨主体类别的强泛化能力。主要难点源于关系中复杂的空间布局、结构变化及微妙的时间动态,导致现有模型过度关注无关视觉细节而忽略有意义的交互。为此,我们提出DreamRelation,一种通过少量示例视频实现关系个性化的新型方法,包含两个核心组件:关系解耦学习与关系动态增强。在关系解耦学习中,采用关系LoRA三元组与混合掩码训练策略,将关系从主体外观中解耦,提升跨关系类型的泛化能力;并通过分析MM-DiT注意力机制中查询、键、值特征的不同作用,确定最优关系LoRA三元组设计,使DreamRelation成为首个具有可解释性的关系视频生成框架。在关系动态增强中,引入时空关系对比损失,强化关系动态建模,降低对主体细节外观的依赖。大量实验表明,DreamRelation在关系视频定制任务上优于当前最先进方法。代码与模型将公开发布。
原文摘要 · Abstract (English)
Relational video customization refers to the creation of personalized videos that depict user-specified relations between two subjects, a crucial task for comprehending real-world visual content. While existing methods can personalize subject appearances and motions, they still struggle with complex relational video customization, where precise relational modeling and high generalization across subject categories are essential. The primary challenge arises from the intricate spatial arrangements, layout variations, and nuanced temporal dynamics inherent in relations; consequently, current models tend to overemphasize irrelevant visual details rather than capturing meaningful interactions. To address these challenges, we propose DreamRelation, a novel approach that personalizes relations through a small set of exemplar videos, leveraging two key components: Relational Decoupling Learning and Relational Dynamics Enhancement. First, in Relational Decoupling Learning, we disentangle relations from subject appearances using relation LoRA triplet and hybrid mask training strategy, ensuring better generalization across diverse relationships. Furthermore, we determine the optimal design of relation LoRA triplet by analyzing the distinct roles of the query, key, and value features within MM-DiT's attention mechanism, making DreamRelation the first relational video generation framework with explainable components. Second, in Relational Dynamics Enhancement, we introduce space-time relational contrastive loss, which prioritizes relational dynamics while minimizing the reliance on detailed subject appearances. Extensive experiments demonstrate that DreamRelation outperforms state-of-the-art methods in relational video customization. Code and models will be made publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。