实现多主体视频的精准动作与身份控制,解决传统方法中身份模糊和退化的难题。
DreamVideo-Omni: Omni-Motion Controlled Multi-Subject Video Customization with Latent Identity Reinforcement Learning
- 分两阶段训练:先融合多维度控制信号,再通过隐空间奖励学习强化身份保持。
- 在多个数据集上实现90%以上的身份保留率和高精度动作控制能力。
- 适合需要精细控制多人动作的视频生成、虚拟人设计等场景。
尽管大规模扩散模型已推动视频生成发展,但对多主体身份与多层次动作的精确控制仍是重大挑战。现有方法常受限于动作粒度不足、控制歧义和身份退化,导致身份保留与动作控制性能不佳。本文提出DreamVideo-Omni,一个统一框架,通过渐进式双阶段训练实现多主体和谐定制与全尺度动作控制。第一阶段整合主体外观、全局运动、局部动态与相机运动等多维控制信号,引入条件感知3D旋转位置编码协调异构输入,并采用分层动作注入策略增强全局运动引导;为解决多主体歧义,设计群体与角色嵌入,显式锚定动作信号至特定身份,有效将复杂场景解耦为独立可控实例。第二阶段为缓解身份退化,基于预训练视频扩散模型设计隐空间身份奖励反馈学习机制,提供动作感知的身份奖励,优先满足人类偏好下的身份保真。依托自建大规模数据集与全面的DreamOmni Bench评估基准,DreamVideo-Omni在生成高质量视频方面表现出卓越的可控性与一致性。
原文摘要 · Abstract (English)
While large-scale diffusion models have revolutionized video synthesis, achieving precise control over both multi-subject identity and multi-granularity motion remains a significant challenge. Recent attempts to bridge this gap often suffer from limited motion granularity, control ambiguity, and identity degradation, leading to suboptimal performance on identity preservation and motion control. In this work, we present DreamVideo-Omni, a unified framework enabling harmonious multi-subject customization with omni-motion control via a progressive two-stage training paradigm. In the first stage, we integrate comprehensive control signals for joint training, encompassing subject appearances, global motion, local dynamics, and camera movements. To ensure robust and precise controllability, we introduce a condition-aware 3D rotary positional embedding to coordinate heterogeneous inputs and a hierarchical motion injection strategy to enhance global motion guidance. Furthermore, to resolve multi-subject ambiguity, we introduce group and role embeddings to explicitly anchor motion signals to specific identities, effectively disentangling complex scenes into independent controllable instances. In the second stage, to mitigate identity degradation, we design a latent identity reward feedback learning paradigm by training a latent identity reward model upon a pretrained video diffusion backbone. This provides motion-aware identity rewards in the latent space, prioritizing identity preservation aligned with human preferences. Supported by our curated large-scale dataset and the comprehensive DreamOmni Bench for multi-subject and omni-motion control evaluation, DreamVideo-Omni demonstrates superior performance in generating high-quality videos with precise controllability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。