arXiv:2509.17506cs.CV2025-09被引 2

4D-MoDe通过分离动静态内容,实现高效可编辑的体积视频流传输。

4D-MoDe: Towards Editable and Scalable Volumetric Streaming via Motion-Decoupled 4D Gaussian Compression

  • 分层表示分离静态背景与动态前景,减少冗余数据。
  • 每帧仅需11.4KB存储,压缩效率提升一个数量级。
  • 支持背景替换、只传前景等实用编辑功能,适合虚拟/增强现实应用。

体积视频已成为沉浸式远程呈现及虚实融合体验的关键媒介,支持六自由度导航与真实空间交互。然而,由于数据量巨大、运动复杂且现有表示方式编辑性差,大规模高质量动态体积内容的传输仍面临挑战。本文提出4D-MoDe,一种基于运动解耦的4D高斯压缩框架,支持可扩展、可编辑的体积视频流。方法采用分层表示,通过前瞻式运动分解策略显式分离静态背景与动态前景,显著降低时间冗余,并支持选择性地流式传输背景或前景。为捕捉连续运动轨迹,引入多分辨率运动估计网格与轻量级共享MLP,并结合动态高斯补偿机制建模新出现内容。自适应分组策略动态插入背景关键帧,平衡时序一致性和压缩效率。此外,设计熵感知训练流程,在率失真目标下联合优化运动场与高斯参数,并采用基于范围和KD树的压缩技术最小化存储开销。在多个数据集上的实验表明,4D-MoDe在保持竞争力重建质量的同时,存储成本较现有最优方法降低一个数量级(如每帧低至11.4 KB),并支持背景替换、前景仅流等实际应用场景。

原文摘要 · Abstract (English)

Volumetric video has emerged as a key medium for immersive telepresence and augmented/virtual reality, enabling six-degrees-of-freedom (6DoF) navigation and realistic spatial interactions. However, delivering high-quality dynamic volumetric content at scale remains challenging due to massive data volume, complex motion, and limited editability of existing representations. In this paper, we present 4D-MoDe, a motion-decoupled 4D Gaussian compression framework designed for scalable and editable volumetric video streaming. Our method introduces a layered representation that explicitly separates static backgrounds from dynamic foregrounds using a lookahead-based motion decomposition strategy, significantly reducing temporal redundancy and enabling selective background/foreground streaming. To capture continuous motion trajectories, we employ a multi-resolution motion estimation grid and a lightweight shared MLP, complemented by a dynamic Gaussian compensation mechanism to model emergent content. An adaptive grouping scheme dynamically inserts background keyframes to balance temporal consistency and compression efficiency. Furthermore, an entropy-aware training pipeline jointly optimizes the motion fields and Gaussian parameters under a rate-distortion (RD) objective, while employing range-based and KD-tree compression to minimize storage overhead. Extensive experiments on multiple datasets demonstrate that 4D-MoDe consistently achieves competitive reconstruction quality with an order of magnitude lower storage cost (e.g., as low as \textbf{11.4} KB/frame) compared to state-of-the-art methods, while supporting practical applications such as background replacement and foreground-only streaming.

体积视频高斯压缩可编辑流运动解耦

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。