用时空哈希编码分离动态与静态特征,提升视频生成质量
STGV: Spatio-Temporal Hash Encoding for Gaussian-based Video Representation

- 将视频特征拆分为2D空间与3D时间哈希编码,解耦动静成分
- 相比现有方法提升0.98 dB PSNR,动态运动更准确
- 适合需要高精度视频重建与动态建模的场景
二维高斯点阵(2DGS)已成为高质量视频表示的有前景范式。然而,现有方法采用内容无关或时空特征重叠嵌入来预测原始高斯基元变形,导致视频中静态与动态成分混杂,难以有效建模其各自特性,进而引发时空变形预测不准和表示质量不佳。为此,本文提出一种基于高斯视频表示的时空哈希编码框架(STGV)。通过将视频特征分解为可学习的2D空间与3D时间哈希编码,STGV有效促进动态成分运动模式的学习,同时保留静态背景细节。此外,我们引入关键帧原始初始化策略,构建更稳定一致的初始高斯表示,避免特征重叠与结构不连贯问题。实验表明,该方法在视频表示质量上优于其他高斯基方法(+0.98 PSNR),并在下游视频任务中达到具有竞争力的性能。
原文摘要 · Abstract (English)
2D Gaussian Splatting (2DGS) has recently become a promising paradigm for high-quality video representation. However, existing methods employ content-agnostic or spatio-temporal feature overlapping embeddings to predict canonical Gaussian primitive deformations, which entangles static and dynamic components in videos and prevents modeling their distinct properties effectively. These result in inaccurate predictions for spatio-temporal deformations and unsatisfactory representation quality. To address these problems, this paper proposes a Spatio-Temporal hash encoding framework for Gaussian-based Video representation (STGV). By decomposing video features into learnable 2D spatial and 3D temporal hash encodings, STGV effectively facilitates the learning of motion patterns for dynamic components while maintaining background details for static elements. In addition, we construct a more stable and consistent initial canonical Gaussian representation through a key frame canonical initialization strategy, preventing from feature overlapping and a structurally incoherent geometry representation. Experimental results demonstrate that our method attains better video representation quality (+0.98 PSNR) against other Gaussian-based methods and achieves competitive performance in downstream video tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。