arXiv:2607.21448cs.CV2026-07

提出新方法实现动态场景高效新视角合成,兼顾精度与速度。

GrainGS: Gradient-Decoupled Gaussian Splatting for Efficient Dynamic Novel View Synthesis

论文配图:GrainGS: Gradient-Decoupled Gaussian Splatting for Efficient Dynamic Novel View Synthesis
图 1 · 摘自论文原文
  • 用分层锚点框架+独立形变,结构稳定且支持局部运动。
  • 合成质量达36.98dB,实时渲染435.6帧/秒,仅需4.67MB存储。
  • 适合追求高精度动态重建的视觉算法研究者。

基于3D高斯泼溅的动态场景重建需在精细运动建模、结构稳定性和紧凑表示间取得平衡。现有逐基元方法虽支持灵活局部变形,但易导致基元冗余增长;锚点法虽提升空间规则性,却抑制了局部运动变化。为此,本文提出GrainGS,融合分层锚点骨架与逐高斯形变的动态高斯框架。首先通过静态预热阶段构建跨时间戳的时不变基准表示;联合训练中,停梯度操作阻断形变对基准位置的梯度传播,同时保留其通过重建目标的直接优化路径。每个高斯独立预测位置、旋转和尺度的时序偏移,实现结构约束下的细节局部运动。此外,采用基准-残差外观分解,建模帧级光度变化而不强制融入几何形变。在合成单目与真实多视角基准上实验表明,GrainGS实现高重建质量、实时新视角合成与紧凑存储:合成基准下平均峰值信噪比达36.98 dB,渲染速度为435.6帧/秒,存储仅需4.67 MB。

原文摘要 · Abstract (English)

Dynamic scene reconstruction with 3D Gaussian Splatting requires a balance between fine-grained motion modeling, structural stability, and compact representation. Existing per-primitive methods provide flexible local deformation but often suffer from redundant primitive growth, while anchor-based methods improve spatial regularity at the cost of suppressing locally varying motion. To address these issues, we present GrainGS, a dynamic Gaussian framework that combines a hierarchical anchor scaffold with per-Gaussian deformation. A static warm-up stage first establishes a time-invariant canonical representation from observations across all timestamps. During joint training, a stop-gradient operation blocks the deformation-mediated gradient pathway to the canonical positions while preserving their direct refinement through the reconstruction objective. Each Gaussian then predicts independent temporal offsets for position, rotation, and scale, enabling detailed local motion within a structurally constrained scaffold. A canonical-residual appearance decomposition further models frame-dependent photometric changes without forcing them into geometric deformation. Experiments on synthetic monocular and real-world multiview benchmarks show that GrainGS achieves high reconstruction quality, real-time novel view synthesis, and compact storage. Under the synthetic benchmark setting, it reaches an average peak signal-to-noise ratio of 36.98 decibels, renders at 435.6 frames per second, and requires 4.67 megabytes of storage.

动态重建高斯泼溅新视角合成实时渲染

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。