通过上下文学习实现高质量3D角色动画,解决动作连贯性与结构保真难题。
SCAIL: Towards Studio-Grade Character Animation via In-Context Learning of 3D-Consistent Pose Representations
- 设计新型3D姿态表示,提供稳定灵活的动作信号。
- 在扩散-变压器中引入全序列姿态注入,提升时空推理能力。
- 构建高质量数据集与评测基准,适合影视级动画研发者使用。
尽管近期取得进展,实现符合影视制作标准的可控角色动画仍具挑战。现有方法虽能将驱动视频中的动作迁移至参考图像,但在复杂运动和跨身份动画等真实场景下,常出现结构失真与时间不一致问题。本文提出SCAIL(面向影视级角色动画的上下文学习框架),通过两项关键创新应对上述挑战:首先,提出一种新型3D姿态表示,提供鲁棒且灵活的动作信号;其次,在扩散-变压器中引入全上下文姿态注入机制,实现对完整动作序列的有效时空推理。为满足影视级要求,我们构建了兼顾多样性和质量的定制化数据流水线,并建立了综合性评测基准。实验表明,SCAIL达到当前最优性能,推动角色动画向影视级控制迈进。代码与模型开源于https://github.com/zai-org/SCAIL。
原文摘要 · Abstract (English)
Achieving controllable character animation that meets studio-grade standards remains challenging despite recent progress. Existing approaches can transfer motion from a driving video to a reference image, but often fail to preserve structural fidelity and temporal consistency in wild scenarios involving complex motion and cross-identity animations. In this work, we present \textbf{SCAIL} (a framework toward \textbf{S}tudio-grade \textbf{C}haracter \textbf{A}nimation via \textbf{I}n-context \textbf{L}earning), which is designed to address these challenges from two key innovations. First, we propose a novel 3D pose representation, providing a robust and flexible motion signal. Second, we introduce a full-context pose injection mechanism within a diffusion-transformer, enabling effective spatio-temporal reasoning over full motion sequences. To align with studio-grade requirements, we develop a curated data pipeline ensuring both diversity and quality, and establish a comprehensive benchmark for systematic evaluation. Experiments show that \textbf{SCAIL} achieves state-of-the-art performance and advances character animation toward studio-grade controlling. Code and model are available at \href{https://github.com/zai-org/SCAIL}{zai-org/SCAIL}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。