小视角视频也能精准还原动态3D场景,突破传统方法局限。
Dynamic View Synthesis from Small Camera Motion Videos
- 用分布正则化替代深度损失,更准确建模场景几何
- 在小运动视频下仍能保持高精度渲染,优于现有方法
- 适合相机移动范围受限的动态场景重建任务
动态3D场景的新视角合成面临重大挑战。尽管基于NeRF的方法取得了显著成果,但其依赖充足的视差信息。当相机运动范围受限甚至静止(即小运动)时,现有方法在场景几何表示和相机参数估计上均出现偏差,导致结果不可靠或失效。为此,我们提出分布型深度正则化(DDR),通过Gumbel-softmax可微采样离散渲染权重分布,计算误差期望,使渲染权重分布与真实分布对齐。同时引入空间点沿射线在物体边界前密度趋近零的约束,确保模型学习正确几何结构。为解释DDR机制,我们设计可视化工具,实现渲染权重层面的几何表征观察。针对相机参数估计问题,我们在训练中引入相机参数学习以增强鲁棒性。大量实验表明,本方法在小相机运动输入下仍能有效重建场景,性能优于当前最优方法。
原文摘要 · Abstract (English)
Novel view synthesis for dynamic $3$D scenes poses a significant challenge. Many notable efforts use NeRF-based approaches to address this task and yield impressive results. However, these methods rely heavily on sufficient motion parallax in the input images or videos. When the camera motion range becomes limited or even stationary (i.e., small camera motion), existing methods encounter two primary challenges: incorrect representation of scene geometry and inaccurate estimation of camera parameters. These challenges make prior methods struggle to produce satisfactory results or even become invalid. To address the first challenge, we propose a novel Distribution-based Depth Regularization (DDR) that ensures the rendering weight distribution to align with the true distribution. Specifically, unlike previous methods that use depth loss to calculate the error of the expectation, we calculate the expectation of the error by using Gumbel-softmax to differentiably sample points from discrete rendering weight distribution. Additionally, we introduce constraints that enforce the volume density of spatial points before the object boundary along the ray to be near zero, ensuring that our model learns the correct geometry of the scene. To demystify the DDR, we further propose a visualization tool that enables observing the scene geometry representation at the rendering weight level. For the second challenge, we incorporate camera parameter learning during training to enhance the robustness of our model to camera parameters. We conduct extensive experiments to demonstrate the effectiveness of our approach in representing scenes with small camera motion input, and our results compare favorably to state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。