arXiv:2505.02192cs.CVcs.AI2025-05ICCV被引 10

让视频生成同时保持身份与动作一致,且不损失任何细节。

DualReal: Adaptive Joint Training for Lossless Identity-Motion Fusion in Video Customization

  • 动态切换训练阶段,分步学习身份与动作,防止信息泄露。
  • 在多个去噪阶段精细控制,实现身份与动作的无损融合。
  • 适合需要高保真定制视频生成的研究者与开发者。

利用预训练大模型进行文本到视频的个性化生成,近年来受到广泛关注,重点在于保持主体身份与动作的一致性。现有方法通常采用孤立定制范式,仅分别定制身份或运动动态,忽略了身份与运动之间的内在约束和协同依赖,导致生成过程中出现身份-动作冲突,系统性地降低质量。为此,我们提出DualReal,一种新颖的自适应联合训练框架,通过协同构建维度间的相互依赖关系。DualReal包含两个模块:(1) 双重感知自适应单元,动态切换训练步骤(即身份或运动),在冻结维度先验指导下学习当前信息,并采用正则化策略避免知识泄露;(2) 阶段融合控制器,利用去噪阶段和扩散变压器深度,以自适应粒度引导不同维度,避免各阶段冲突,最终实现身份与动作模式的无损融合。我们构建了比现有方法更全面的评估基准。实验结果表明,DualReal在CLIP-I和DINO-I指标上平均提升21.7%和31.8%,并在几乎所有运动指标上达到顶尖表现。

原文摘要 · Abstract (English)

Customized text-to-video generation with pre-trained large-scale models has recently garnered significant attention by focusing on identity and motion consistency. Existing works typically follow the isolated customized paradigm, where the subject identity or motion dynamics are customized exclusively. However, this paradigm completely ignores the intrinsic mutual constraints and synergistic interdependencies between identity and motion, resulting in identity-motion conflicts throughout the generation process that systematically degrade. To address this, we introduce DualReal, a novel framework that employs adaptive joint training to construct interdependencies between dimensions collaboratively. Specifically, DualReal is composed of two units: (1) Dual-aware Adaptation dynamically switches the training step (i.e., identity or motion), learns the current information guided by the frozen dimension prior, and employs a regularization strategy to avoid knowledge leakage; (2) StageBlender Controller leverages the denoising stages and Diffusion Transformer depths to guide different dimensions with adaptive granularity, avoiding conflicts at various stages and ultimately achieving lossless fusion of identity and motion patterns. We constructed a more comprehensive evaluation benchmark than existing methods. The experimental results show that DualReal improves CLIP-I and DINO-I metrics by 21.7% and 31.8% on average, and achieves top performance on nearly all motion metrics. Page: https://wenc-k.github.io/dualreal-customization

视频生成身份一致动作融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。