用3D轨迹精准控制骨骼动画,提升长时一致性与细节精度。
Follow Your Track: Precise Skeleton Animation Controlled by 3D Trajectories

- 以骨骼为结构表示,用单目视频轨迹做运动引导,避免外观干扰。
- 在多个数据集上实现更高保真度和更优时间连续性,最长生成10秒动画。
- 适合需要精细动作控制的虚拟人、角色动画等场景使用。
4D生成旨在为3D物体添加真实运动,前景广阔。现有方法通常将3D资产生成与运动合成分离:先获取3D资产,构建网格或高斯表示,再通过文本或视频控制信号生成运动。但密集网格和高斯表示计算开销大,易产生时间伪影,限制了动画质量与长度,仅能生成短片段。同时,文本缺乏时空细节(如时机、协调性),视频则将运动与外观、背景纠缠在一起。这些限制导致4D动画存在时间不一致、识别错误、可控性差等问题。我们提出 exttt{ACT},一种基于轨迹的通用骨骼动画框架。ACT采用骨骼作为紧凑且高效的结构表示,并利用单目视频提取的3D点轨迹作为显式运动引导,提供详细运动模式而不含外观信息。核心是路由轨迹注入器,通过三种互补设计实现精确可靠的轨迹到关节映射:先验引导的硬路由建立精确的骨架-网格对应关系,全局路由实现全身关节-轨迹交互以增强整体运动感知,局部窗口交叉注意力强化细微时间对齐,改善微调时机并减少不同运动速率下的错位。大量实验表明, exttt{ACT} 在保真度和时间一致性方面显著优于现有方法。
原文摘要 · Abstract (English)
4D generation aims to animate 3D objects with realistic motion, holding great promise for applications. Existing methods typically decouple 3D asset generation from motion synthesis: acquire a 3D asset, prepare a structural representation like mesh and Gaussians, and synthesize motion from text or video control signals. However, dense mesh and Gaussian representations incur high computational costs and are prone to temporal artifacts, limiting animation quality and duration to only short clips. Meanwhile, text lacks fine-grained spatial and temporal details such as timing and coordination, while video entangles motion with appearance and background. Together, these limitations result in 4D animations that suffer from poor temporal consistency, wrong identification, and limited controllability. We address these issues with \texttt{ACT}, a trajectory-conditioned framework for topology-general skeletal animation. ACT uses skeletons as a compact structured and compute-efficient representation and 3D point trajectories from monocular video as explicit motion guidance which provide detailed motion patterns without appearance entanglement. At the core of ACT is a Routed Trajectory Injector, which achieves accurate and robust trajectory-to-joint transfer through three complementary designs: prior-guided hard routing establishes precise skeleton-to-mesh correspondences, global routing enables holistic joint-track interaction for full-body motion awareness, and local windowed cross-attention enforces fine-grained temporal alignment, improving micro-timing and reducing motion misalignment across varying motion rates. Extensive experiments demonstrate that \texttt{ACT} significantly outperforms existing methods in fidelity and temporal consistency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。