用精确深度信息提升神经3D场景重建的几何与光照还原效果
Incorporating dense metric depth into neural 3D representations for view synthesis and relighting
- 将密集度量深度融入神经3D表示训练,优化几何估计
- 在仅有少数视角下实现高质量视图合成与可控光照重渲染
- 适用于机器人、虚拟现实等受限视角场景
在游戏、虚拟现实、机器人操作、自动驾驶及消费级摄影等领域,准确重建小场景的几何与逼真外观是研究热点。然而,在机器人应用中,由于运动范围有限和场景遮挡,现有方法常因视角稀疏导致重建质量差甚至失败。而实际中,可通过立体视觉直接获取密集度量深度,且光照可控制。本文提出一种将密集度量深度融入神经3D表示训练的方法,并通过区分纹理与几何边缘解决联合优化中的伪影问题。同时设计了一套多闪光立体相机系统以采集所需数据,展示了仅需少量训练视图即可实现高质量视图合成与光照重渲染的效果。
原文摘要 · Abstract (English)
Synthesizing accurate geometry and photo-realistic appearance of small scenes is an active area of research with compelling use cases in gaming, virtual reality, robotic-manipulation, autonomous driving, convenient product capture, and consumer-level photography. When applying scene geometry and appearance estimation techniques to robotics, we found that the narrow cone of possible viewpoints due to the limited range of robot motion and scene clutter caused current estimation techniques to produce poor quality estimates or even fail. On the other hand, in robotic applications, dense metric depth can often be measured directly using stereo and illumination can be controlled. Depth can provide a good initial estimate of the object geometry to improve reconstruction, while multi-illumination images can facilitate relighting. In this work we demonstrate a method to incorporate dense metric depth into the training of neural 3D representations and address an artifact observed while jointly refining geometry and appearance by disambiguating between texture and geometry edges. We also discuss a multi-flash stereo camera system developed to capture the necessary data for our pipeline and show results on relighting and view synthesis with a few training views.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。