让游戏角色动起来:支持动态背景的通用图像动画框架
Animate-X++: Universal Character Image Animation with Dynamic Backgrounds
- 用运动指示器融合隐式与显式动作特征,提升跨角色泛化能力
- 联合训练动画与文本生成背景,实现角色与环境同步动态
- 新基准测试验证对非人类角色的通用性,性能显著优于现有方法
角色图像动画近年来取得显著进展,可从参考图像和目标姿态序列生成高质量视频。然而,多数方法仅适用于人体,难以推广到游戏与娱乐产业中常见的拟人化角色。此外,以往方法仅能生成静态背景视频,限制了真实感。本文分析发现,问题根源在于运动建模不足,无法理解驱动视频的动作模式,导致姿态序列被僵硬套用。为此,提出基于DiT的通用动画框架Animate-X++,引入运动指示器(Pose Indicator),通过CLIP视觉特征提取驱动视频的总体运动模式与时间关系(隐式),并预先模拟推理可能输入以增强模型泛化性(显式)。针对背景静态问题,采用多任务训练策略,联合优化动画与文本到视频(TI2V)任务,并结合部分参数训练,实现角色动画与文本驱动的动态背景生成。构建新基准A2Bench评估泛化能力。大量实验表明,Animate-X++在各类角色上均表现优异,显著超越现有方法。
原文摘要 · Abstract (English)
Character image animation, which generates high-quality videos from a reference image and target pose sequence, has seen significant progress in recent years. However, most existing methods only apply to human figures, which usually do not generalize well on anthropomorphic characters commonly used in industries like gaming and entertainment. Furthermore, previous methods could only generate videos with static backgrounds, which limits the realism of the videos. For the first challenge, our in-depth analysis suggests to attribute this limitation to their insufficient modeling of motion, which is unable to comprehend the movement pattern of the driving video, thus imposing a pose sequence rigidly onto the target character. To this end, this paper proposes Animate-X++, a universal animation framework based on DiT for various character types, including anthropomorphic characters. To enhance motion representation, we introduce the Pose Indicator, which captures comprehensive motion pattern from the driving video through both implicit and explicit manner. The former leverages CLIP visual features of a driving video to extract its gist of motion, like the overall movement pattern and temporal relations among motions, while the latter strengthens the generalization of DiT by simulating possible inputs in advance that may arise during inference. For the second challenge, we introduce a multi-task training strategy that jointly trains the animation and TI2V tasks. Combined with the proposed partial parameter training, this approach achieves not only character animation but also text-driven background dynamics, making the videos more realistic. Moreover, we introduce a new Animated Anthropomorphic Benchmark (A2Bench) to evaluate the performance of Animate-X++ on universal and widely applicable animation images. Extensive experiments demonstrate the superiority and effectiveness of Animate-X++.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。