arXiv:2409.08353cs.GRcs.CV2024-09中稿 · SIGGRAPH被引 48

用双高斯分解人体动作与外观,实现超高压缩的沉浸式人像视频。

Robust Dual Gaussian Splatting for Immersive Human-centric Volumetric Videos

  • 分离皮肤与关节高斯表示,显式解耦动作与外观信息。
  • 压缩比高达120倍,每帧仅需约350KB存储空间。
  • 适合虚拟现实中的高保真人体视频实时播放场景。

体积分层视频代表了视觉媒体的革命性进步,使用户能够自由漫游沉浸式虚拟体验,缩小数字世界与真实世界的差距。然而,现有流程需要大量人工干预来稳定网格序列,并生成过于庞大的资产,阻碍了广泛应用。本文提出一种基于高斯的新方法——DualGS,用于实现实时、高保真的人体表演播放,具备优异的压缩比。DualGS的核心思想是分别用皮肤高斯和关节高斯表示运动与外观。这种显式解耦能显著减少运动冗余并提升时间一致性。我们首先在首帧将皮肤高斯锚定到关节高斯,随后采用粗到精的训练策略进行逐帧人体表演建模,包括整体运动预测的粗对齐阶段和鲁棒跟踪与高保真渲染的细粒度优化。为将体积分层视频无缝集成至VR环境,我们使用熵编码高效压缩运动,用编解码器压缩外观并结合持久代码本。该方法实现了最高120倍的压缩比,每帧仅需约350KB存储空间。我们在VR头显上展示了该表示的有效性,实现了照片级真实的自由视角体验,让用户沉浸式观看音乐家表演,感受指尖传递的节奏。

原文摘要 · Abstract (English)

Volumetric video represents a transformative advancement in visual media, enabling users to freely navigate immersive virtual experiences and narrowing the gap between digital and real worlds. However, the need for extensive manual intervention to stabilize mesh sequences and the generation of excessively large assets in existing workflows impedes broader adoption. In this paper, we present a novel Gaussian-based approach, dubbed \textit{DualGS}, for real-time and high-fidelity playback of complex human performance with excellent compression ratios. Our key idea in DualGS is to separately represent motion and appearance using the corresponding skin and joint Gaussians. Such an explicit disentanglement can significantly reduce motion redundancy and enhance temporal coherence. We begin by initializing the DualGS and anchoring skin Gaussians to joint Gaussians at the first frame. Subsequently, we employ a coarse-to-fine training strategy for frame-by-frame human performance modeling. It includes a coarse alignment phase for overall motion prediction as well as a fine-grained optimization for robust tracking and high-fidelity rendering. To integrate volumetric video seamlessly into VR environments, we efficiently compress motion using entropy encoding and appearance using codec compression coupled with a persistent codebook. Our approach achieves a compression ratio of up to 120 times, only requiring approximately 350KB of storage per frame. We demonstrate the efficacy of our representation through photo-realistic, free-view experiences on VR headsets, enabling users to immersively watch musicians in performance and feel the rhythm of the notes at the performers' fingertips.

体积分层高斯溅射压缩虚拟现实

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。