arXiv:2504.14977cs.CV2025-04被引 30

用简单修改+高效微调,让视频模型更可控地生成真实场景角色动画。

RealisDance-DiT: Simple yet Strong Baseline towards Controllable Character Animation in the Wild

  • 在大模型基础上做最小改动,配合特殊微调策略提升可控性。
  • 新数据集下性能超越现有方法,尤其在复杂光照和互动场景中表现突出。
  • 适合研究可控视频生成、角色动画的开发者参考使用。

可控角色动画仍是难题,尤其在罕见姿态、风格化角色、人物与物体交互、复杂光照及动态场景中。以往工作多通过复杂的旁路网络注入姿态和外观引导,但难以泛化到开放世界。本文提出新思路:只要基础模型足够强大,仅需简单修改搭配灵活微调策略,即可有效应对上述挑战,迈向真实世界中的可控角色动画。我们基于 Wan-2.1 视频基础模型构建 RealisDance-DiT。充分分析表明,广泛采用的 Reference Net 设计对大规模 DiT 模型并不最优。相反,我们证明对基础模型架构进行最小调整即可形成强劲基线。此外,提出低噪声预热和“大批次小迭代”微调策略,加速收敛同时最大程度保留基础模型先验。还引入一个新测试数据集,涵盖多样真实世界挑战,补充 TikTok 数据集和 UBC 时尚视频数据集,全面评估方法。大量实验显示,RealisDance-DiT 显著优于现有方法。

原文摘要 · Abstract (English)

Controllable character animation remains a challenging problem, particularly in handling rare poses, stylized characters, character-object interactions, complex illumination, and dynamic scenes. To tackle these issues, prior work has largely focused on injecting pose and appearance guidance via elaborate bypass networks, but often struggles to generalize to open-world scenarios. In this paper, we propose a new perspective that, as long as the foundation model is powerful enough, straightforward model modifications with flexible fine-tuning strategies can largely address the above challenges, taking a step towards controllable character animation in the wild. Specifically, we introduce RealisDance-DiT, built upon the Wan-2.1 video foundation model. Our sufficient analysis reveals that the widely adopted Reference Net design is suboptimal for large-scale DiT models. Instead, we demonstrate that minimal modifications to the foundation model architecture yield a surprisingly strong baseline. We further propose the low-noise warmup and "large batches and small iterations" strategies to accelerate model convergence during fine-tuning while maximally preserving the priors of the foundation model. In addition, we introduce a new test dataset that captures diverse real-world challenges, complementing existing benchmarks such as TikTok dataset and UBC fashion video dataset, to comprehensively evaluate the proposed method. Extensive experiments show that RealisDance-DiT outperforms existing methods by a large margin.

角色动画视频生成微调策略可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。