解决长时序动态场景4D高斯渲染的闪烁与内存爆炸问题
MoRel: Long-Range Flicker-Free 4D Motion Modeling via Anchor Relay-based Bidirectional Blending with Hierarchical Densification
- 通过锚点中继双向融合机制建模帧间形变,提升时间一致性
- 在自建长时序数据集上实现无闪烁、低内存的4D重建
- 适合需要高效处理长时间动态视频的科研与工业应用
近年来,4D高斯点阵(4DGS)将3D高斯点阵(3DGS)的高速渲染能力拓展至时间维度,实现了动态场景的实时渲染。然而,现有方法在建模长时序动态视频时仍面临严重内存爆炸、时间闪烁以及难以处理随时间出现或消失的遮挡等问题。为此,本文提出一种新型4DGS框架MoRel,其核心为基于锚点中继的双向融合(ARBB)机制,可在关键帧处逐步构建局部规范锚点空间,并在锚点层面建模帧间形变,增强时间连贯性。通过学习双向形变并结合可学习不透明度控制进行自适应融合,有效缓解了时间不连续和闪烁伪影。此外,提出基于特征方差的分层稠密化(FHD)策略,在保持渲染质量的前提下对锚点空间进行有效稠密化。为评估模型在真实长时序4D运动下的表现,我们构建了新的长时序4D运动数据集SelfCap$_{\text{LR}}$,其平均动态运动幅度更大,覆盖空间更广。实验表明,MoRel在保持有界内存消耗的同时,实现了时间一致且无闪烁的长时序4D重建,展现出良好的可扩展性与效率。
原文摘要 · Abstract (English)
Recent advances in 4D Gaussian Splatting (4DGS) have extended the high-speed rendering capability of 3D Gaussian Splatting (3DGS) into the temporal domain, enabling real-time rendering of dynamic scenes. However, one of the major remaining challenges lies in modeling long-range motion-contained dynamic videos, where a naive extension of existing methods leads to severe memory explosion, temporal flickering, and failure to handle appearing or disappearing occlusions over time. To address these challenges, we propose a novel 4DGS framework characterized by an Anchor Relay-based Bidirectional Blending (ARBB) mechanism, named MoRel, which enables temporally consistent and memory-efficient modeling of long-range dynamic scenes. Our method progressively constructs locally canonical anchor spaces at key-frame time index and models inter-frame deformations at the anchor level, enhancing temporal coherence. By learning bidirectional deformations between KfA and adaptively blending them through learnable opacity control, our approach mitigates temporal discontinuities and flickering artifacts. We further introduce a Feature-variance-guided Hierarchical Densification (FHD) scheme that effectively densifies KfA's while keeping rendering quality, based on an assigned level of feature-variance. To effectively evaluate our model's capability to handle real-world long-range 4D motion, we newly compose long-range 4D motion-contained dataset, called SelfCap$_{\text{LR}}$. It has larger average dynamic motion magnitude, captured at spatially wider spaces, compared to previous dynamic video datasets. Overall, our MoRel achieves temporally coherent and flicker-free long-range 4D reconstruction while maintaining bounded memory usage, demonstrating both scalability and efficiency in dynamic Gaussian-based representations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。