arXiv:2503.12242cs.CV2025-03CVPR被引 13

用高斯表示统一播放与重演,实现逼真人体为中心的动态视频重演。

RePerformer: Immersive Human-centric Volumetric Videos from Playback to Photoreal Reperformance

  • 分层解耦动态场景为运动与外观高斯,在规范空间中关联建模。
  • 用默顿编码高效压缩外观高斯为2D位置与属性图,支持高保真渲染。
  • 通过语义对齐与形变迁移,可实现新动作下的照片级重演,适合虚拟表演应用。

人体为中心的体素视频提供沉浸式自由视角体验,但现有方法要么聚焦于一般动态场景的回放,要么专注于人物角色动画,难以实现一般动态场景的重演。本文提出RePerformer,一种基于高斯的新表示方法,统一了高保真人体为中心体素视频的播放与重演。具体地,我们层次化地将动态场景解耦为运动高斯和外观高斯,并在规范空间中建立关联。进一步采用基于默顿(Morton)的参数化方法,高效将外观高斯编码为二维位置与属性图。为增强泛化能力,我们使用2D CNN将位置图映射为属性图,可组装成外观高斯以实现高保真动态场景渲染。对于重演,我们设计语义感知对齐模块,对运动高斯应用形变迁移,从而在新动作下实现照片级渲染。大量实验验证了RePerformer的鲁棒性与有效性,为人体为中心体素视频的播放-重演范式树立了新基准。

原文摘要 · Abstract (English)

Human-centric volumetric videos offer immersive free-viewpoint experiences, yet existing methods focus either on replaying general dynamic scenes or animating human avatars, limiting their ability to re-perform general dynamic scenes. In this paper, we present RePerformer, a novel Gaussian-based representation that unifies playback and re-performance for high-fidelity human-centric volumetric videos. Specifically, we hierarchically disentangle the dynamic scenes into motion Gaussians and appearance Gaussians which are associated in the canonical space. We further employ a Morton-based parameterization to efficiently encode the appearance Gaussians into 2D position and attribute maps. For enhanced generalization, we adopt 2D CNNs to map position maps to attribute maps, which can be assembled into appearance Gaussians for high-fidelity rendering of the dynamic scenes. For re-performance, we develop a semantic-aware alignment module and apply deformation transfer on motion Gaussians, enabling photo-real rendering under novel motions. Extensive experiments validate the robustness and effectiveness of RePerformer, setting a new benchmark for playback-then-reperformance paradigm in human-centric volumetric videos.

体素视频高斯表示重演生成自由视角

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。