arXiv:2508.07409cs.CV2025-08被引 12

仅用一张图和动作序列,就能生成稳定可控的3D角色动画。

CharacterShot: Controllable and Consistent 4D Character Animation

  • 基于2D图像转视频模型,用动作序列控制角色动画。
  • 通过多视角一致性优化,生成连续稳定的4D角色表示。
  • 适合角色设计、动画制作人员快速生成高质量动画。

本文提出CharacterShot,一种可控且一致的4D角色动画框架,使任意设计师仅需单张参考角色图像和2D动作序列即可创建动态3D角色(即4D动画)。首先,基于先进的DiT架构预训练2D角色动画模型,支持任意2D动作序列作为控制信号。随后,引入双注意力模块与相机先验,将2D模型扩展至3D,生成具有时空与视点一致性的多视角视频。最后,采用新型邻域约束4D高斯点云优化方法,从多视角视频中恢复出连续稳定的4D角色表示。为提升角色中心性能,构建大规模数据集Character4D,包含13,115个不同外观与动作的唯一角色,从多视角渲染而成。在新构建的基准CharacterBench上,大量实验表明本方法优于当前最先进方法。代码、模型与数据集将公开于https://github.com/Jeoyal/CharacterShot。

原文摘要 · Abstract (English)

In this paper, we propose \textbf{CharacterShot}, a controllable and consistent 4D character animation framework that enables any individual designer to create dynamic 3D characters (i.e., 4D character animation) from a single reference character image and a 2D pose sequence. We begin by pretraining a powerful 2D character animation model based on a cutting-edge DiT-based image-to-video model, which allows for any 2D pose sequnce as controllable signal. We then lift the animation model from 2D to 3D through introducing dual-attention module together with camera prior to generate multi-view videos with spatial-temporal and spatial-view consistency. Finally, we employ a novel neighbor-constrained 4D gaussian splatting optimization on these multi-view videos, resulting in continuous and stable 4D character representations. Moreover, to improve character-centric performance, we construct a large-scale dataset Character4D, containing 13,115 unique characters with diverse appearances and motions, rendered from multiple viewpoints. Extensive experiments on our newly constructed benchmark, CharacterBench, demonstrate that our approach outperforms current state-of-the-art methods. Code, models, and datasets will be publicly available at https://github.com/Jeoyal/CharacterShot.

4D动画角色生成可控生成多视角一致

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。