arXiv:2511.09827cs.CV2025-11被引 4

用高斯点云实现人像在复杂场景中自然动画,支持自由视角渲染。

AHA! Animating Human Avatars in Diverse Scenes with Gaussian Splatting

  • 以3D高斯点云表示人与场景,解耦渲染与动作生成
  • 无需成对数据,通过透明度和投影结构引导姿态对齐
  • 可处理单目视频动画,适合影视与虚拟人应用

我们提出一种基于3D高斯点云(3DGS)的新框架,用于在3D场景中动画化人物。3DGS是近期达到顶尖真实感新视角合成效果的神经场景表示方法,但在人-场景动画与交互方面仍待探索。不同于使用网格或点云的传统动画流程,本方法将人物与场景统一表示为高斯点,实现几何一致的自由视角渲染。核心思想是将渲染与运动合成解耦,可独立处理且无需配对的人-场景数据。关键在于一个高斯对齐的运动模块,通过透明度线索与投影高斯结构引导人物定位与姿态对齐,无需显式场景几何。为进一步实现自然交互,提出人体-场景高斯优化策略,强制实现真实接触与导航。我们在Scannet++和SuperSplat数据集上评估,使用稀疏与密集多视角人体采集重建的化身。结果表明,该框架可实现单目RGB视频编辑后的人物动画自由视角渲染,充分展现3DGS在单目视频驱动人物动画中的独特优势。完整结果请参见补充材料:https://miraymen.github.io/aha/

原文摘要 · Abstract (English)

We present a novel framework for animating humans in 3D scenes using 3D Gaussian Splatting (3DGS), a neural scene representation that has recently achieved state-of-the-art photorealistic results for novel-view synthesis but remains under-explored for human-scene animation and interaction. Unlike existing animation pipelines that use meshes or point clouds as the underlying 3D representation, our approach introduces the use of 3DGS as the 3D representation for animating humans in scenes. By representing humans and scenes as Gaussians, our approach allows geometry-consistent free-viewpoint rendering of humans interacting with 3D scenes. Our key insight is that rendering can be decoupled from motion synthesis, and each sub-problem can be addressed independently without the need for paired human-scene data. Central to our method is a Gaussian-aligned motion module that synthesizes motion without explicit scene geometry, using opacity-based cues and projected Gaussian structures to guide human placement and pose alignment. To ensure natural interactions, we further propose a human-scene Gaussian refinement optimization that enforces realistic contact and navigation. We evaluate our approach on scenes from Scannet++ and the SuperSplat library, and on avatars reconstructed from sparse and dense multi-view human capture. Finally, we demonstrate that our framework enables novel applications such as geometry-consistent free-viewpoint rendering of edited monocular RGB videos with newly animated humans, showcasing the unique advantages of 3DGS for monocular video-based human animation. To assess the full quality of our results, we encourage readers to view the supplementary material available at https://miraymen.github.io/aha/ .

3D动画高斯点云场景交互单目视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。