arXiv:2503.05162cs.CV2025-03被引 4

用可演化高斯表示实现高质量动态视频流传输

EvolvingGS: High-Fidelity Streamable Volumetric Video via Evolving 3D Gaussian Representation

  • 分两阶段演化高斯模型,先粗对齐再局部精修
  • 在复杂人体动作序列上实现超50倍压缩率
  • 适合长时序动态场景重建与实时流式播放

近期基于显式点的3D高斯喷溅(3DGS)在3D场景重建中取得显著进展,以高质量和快速渲染著称。然而,对复杂人体表演等长时间动态场景的重建仍具挑战。现有方法难以建模剧烈运动、频繁拓扑变化或与道具交互的长序列,通常将序列分割为独立帧组处理,破坏了时间稳定性,导致观感不佳且存储效率低。为此,我们提出EvolvingGS,一种两阶段策略:首先形变高斯模型粗略对齐目标帧,再通过极少量点增删在快速变化区域进行精修。得益于增量演化表示的灵活性,该方法在逐帧和时间质量指标上均优于现有方法,同时保持纯显式表示带来的快速渲染。此外,利用相邻帧间的时序一致性,提出一种简单高效的压缩算法,实现超过50倍压缩率。在公开基准和挑战性自定义数据集上的大量实验表明,该方法显著推进了长序列复杂动态场景重建的前沿水平。

原文摘要 · Abstract (English)

We have recently seen great progress in 3D scene reconstruction through explicit point-based 3D Gaussian Splatting (3DGS), notable for its high quality and fast rendering speed. However, reconstructing dynamic scenes such as complex human performances with long durations remains challenging. Prior efforts fall short of modeling a long-term sequence with drastic motions, frequent topology changes or interactions with props, and resort to segmenting the whole sequence into groups of frames that are processed independently, which undermines temporal stability and thereby leads to an unpleasant viewing experience and inefficient storage footprint. In view of this, we introduce EvolvingGS, a two-stage strategy that first deforms the Gaussian model to coarsely align with the target frame, and then refines it with minimal point addition/subtraction, particularly in fast-changing areas. Owing to the flexibility of the incrementally evolving representation, our method outperforms existing approaches in terms of both per-frame and temporal quality metrics while maintaining fast rendering through its purely explicit representation. Moreover, by exploiting temporal coherence between successive frames, we propose a simple yet effective compression algorithm that achieves over 50x compression rate. Extensive experiments on both public benchmarks and challenging custom datasets demonstrate that our method significantly advances the state-of-the-art in dynamic scene reconstruction, particularly for extended sequences with complex human performances.

3D重建动态视频高斯表示流式传输

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。