用分层残差学习提升动态场景4D高斯点云渲染质量
CTRL-GS: Cascaded Temporal Residue Learning for 4D Gaussian Splatting
- 将动态场景分解为视频段-帧结构,用光流自适应调整
- 通过常量+残差方式建模时变信号,提升复杂场景表现
- 适合处理大运动、遮挡、细节丰富的动态场景
近期,高斯点云方法已成为多视角图像或视频捕获场景中新视图合成的优选替代方案,优于传统的辐射场方法。本文提出一种面向动态场景的4D高斯点云新扩展。受残差学习启发,我们层次化地将动态场景分解为“视频段-帧”结构,其中视频段通过光流动态调整。不直接预测时变信号,而是将其建模为视频常量、段常量与帧级残差之和。该方法使模型更灵活,能适应高度变化的场景。我们在多个标准数据集上实现了领先视觉质量与实时渲染性能,尤其在具有大运动、遮挡和精细细节的复杂场景中改进最为显著,而当前方法在此类场景下性能下降明显。
原文摘要 · Abstract (English)
Recently, Gaussian Splatting methods have emerged as a desirable substitute for prior Radiance Field methods for novel-view synthesis of scenes captured with multi-view images or videos. In this work, we propose a novel extension to 4D Gaussian Splatting for dynamic scenes. Drawing on ideas from residual learning, we hierarchically decompose the dynamic scene into a "video-segment-frame" structure, with segments dynamically adjusted by optical flow. Then, instead of directly predicting the time-dependent signals, we model the signal as the sum of video-constant values, segment-constant values, and frame-specific residuals, as inspired by the success of residual learning. This approach allows more flexible models that adapt to highly variable scenes. We demonstrate state-of-the-art visual quality and real-time rendering on several established datasets, with the greatest improvements on complex scenes with large movements, occlusions, and fine details, where current methods degrade most.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。