arXiv:2602.23040cs.CV2026-02被引 2

用压缩的纹理图表示4D体视频,支持高效存储与播放。

PackUV: Packed Gaussian UV Maps for 4D Volumetric Video

  • 将高斯点属性映射到多尺度纹理图,实现紧凑图像式存储。
  • 在20亿帧数据上验证,可稳定渲染长达30分钟的视频。
  • 兼容主流视频编码器,适合实际流媒体部署。

体视频提供沉浸式4D体验,但重建、存储和流传输仍面临挑战。现有基于高斯溅射的方法虽能实现高质量重建,但在长序列、时间不一致、大运动和遮挡情况下表现不佳,且输出难以适配传统视频编码流程。本文提出PackUV,一种新型4D高斯表示,将所有高斯属性映射到一系列结构化、多尺度的UV图集,实现紧凑且图像原生的存储。为从多视角视频中拟合该表示,我们提出PackUV-GS,一种在UV域直接优化高斯参数的时序一致性方法。通过光流引导的高斯标记与视频关键帧模块,识别动态高斯点,稳定静态区域,在大运动和遮挡下仍保持时间连贯性。生成的UV图集是首个兼容标准视频编码器(如FFV1)且不损失质量的体视频表示,支持在现有多媒体基础设施中高效流传输。为评估长时体视频捕获,我们构建了PackUV-2B,目前最大多视角视频数据集,包含超过50台同步相机、显著运动与频繁遮挡,覆盖100个序列共20亿帧。大量实验表明,本方法在渲染保真度上超越现有基线,且可扩展至长达30分钟的序列并保持一致质量。

原文摘要 · Abstract (English)

Volumetric videos offer immersive 4D experiences, but remain difficult to reconstruct, store, and stream at scale. Existing Gaussian Splatting based methods achieve high-quality reconstruction but break down on long sequences, temporal inconsistency, and fail under large motions and disocclusions. Moreover, their outputs are typically incompatible with conventional video coding pipelines, preventing practical applications. We introduce PackUV, a novel 4D Gaussian representation that maps all Gaussian attributes into a sequence of structured, multi-scale UV atlas, enabling compact, image-native storage. To fit this representation from multi-view videos, we propose PackUV-GS, a temporally consistent fitting method that directly optimizes Gaussian parameters in the UV domain. A flow-guided Gaussian labeling and video keyframing module identifies dynamic Gaussians, stabilizes static regions, and preserves temporal coherence even under large motions and disocclusions. The resulting UV atlas format is the first unified volumetric video representation compatible with standard video codecs (e.g., FFV1) without losing quality, enabling efficient streaming within existing multimedia infrastructure. To evaluate long-duration volumetric capture, we present PackUV-2B, the largest multi-view video dataset to date, featuring more than 50 synchronized cameras, substantial motion, and frequent disocclusions across 100 sequences and 2B (billion) frames. Extensive experiments demonstrate that our method surpasses existing baselines in rendering fidelity while scaling to sequences up to 30 minutes with consistent quality.

4D视频高斯溅射视频编码体渲染

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。