arXiv:2608.28219cs.CV2026-08中稿 · ECCV

分离空间与运动先验,让动画角色更自然地跟随动作。

RASA: Disentangled Spatial-Motional Priors for Cross-Identity Character Animation

论文配图:RASA: Disentangled Spatial-Motional Priors for Cross-Identity Character Animation
图 1 · 摘自论文原文
  • 用结构化先验分离角色位置与动作控制
  • 在高保真数据集上显著提升动作真实度和视觉质量
  • 适合需要精细角色动画的研究者与开发者

跨身份角色动画旨在从参考图像生成目标身份,并使其动作模仿驱动视频中的源角色。核心挑战在于空间映射(对齐位置、尺度和骨骼比例)与运动控制(优化关节运动、体积一致性和视角一致性)的固有耦合。我们提出参考感知结构对齐(RASA),通过向扩散Transformer(DiT)注入结构化先验,将空间映射与运动控制解耦。该方法分两阶段:首先,空间先验校准器(SPC)融合参考身份与驱动姿态,生成空间对齐的初始噪声隐变量,确保正确的位置、尺度及与驱动骨架的对齐;其次,内在运动引导器(IMG)将与形状无关的SMPL运动参数编码为语义运动向量,超越外观依赖的2D关键点。该向量注入DiT中间层,补充基础姿态条件,实现解剖学一致的关节运动与视角感知的体积优化。我们构建了CIM-Bench,一个高质量、严格筛选的基准数据集用于评估。大量实验表明,RASA在动作保真度与视觉质量上显著优于现有最先进方法。本工作确立新范式:解耦的空间与运动先验是实现鲁棒角色动画的关键。

原文摘要 · Abstract (English)

Cross-identity character animation aims to drive a target identity from a reference image to follow the motion of a source character from a driving video. The core challenge lies in the inherent entanglement of two capabilities: cross-identity spatial mapping (aligning position, scale, and skeletal proportions) and motion control (refining joint articulation, volumetric consistency, and view coherence). We introduce Reference-Aware Structural Alignment (RASA), a framework that disentangles spatial mapping from motion control by injecting structured priors into a Diffusion Transformer (DiT). Our approach has two stages. First, a Spatial Prior Calibrator (SPC) fuses reference identity with driving pose to generate a spatially grounded initial noise latent, ensuring correct positioning, scaling, and alignment with the driving skeleton. Second, an Inherent Motional Guider (IMG) encodes shape-agnostic SMPL articulation parameters into a semantic motion vector beyond appearance-biased 2D keypoints. Injected into intermediate DiT layers, this vector complements the base pose condition for anatomically consistent articulation and view-aware volumetric refinement. We curate CIM-Bench, a high-quality benchmark with rigorous curation, for evaluation. Extensive experiments show RASA significantly outperforms state-of-the-art methods in motion fidelity and visual quality. Our work establishes a new paradigm showing disentangled spatial and motional priors are key to robust character animation. Project page: https://hidream.ai.github.io/RASA/

角色动画扩散模型解耦表示运动控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。