arXiv:2409.17924cs.CV2024-09SIGGRAPH被引 8

用神经光球模型实现手机全景拼接与重渲染,支持动态场景和任意路径拍摄。

Neural Light Spheres for Implicit Image Stitching and View Synthesis

  • 用单层神经光球建模,实时估计相机路径和高分辨率场景
  • 80MB模型大小,1080p下50帧/秒实时渲染,支持视点依赖光照
  • 适合移动设备全景拍摄,尤其抗运动模糊和非理想拍摄条件

捕捉与显示均具挑战性,全景图在现代手机相机中虽常见却未被充分利用。本文提出一种球形神经光场模型,用于隐式全景图像拼接与重渲染,可处理深度视差、视点相关光照及拍摄过程中的局部场景运动与颜色变化。该模型在测试时拟合任意路径的全景视频(垂直、水平、随机行走),联合估计相机轨迹与高分辨率场景重建,生成宽视场的新视角投影。单层结构避免了昂贵的体素采样,将场景分解为紧凑的视点相关射线偏移与颜色分量,每场景模型仅80MB,1080p下实现50帧/秒实时渲染。实验表明,其重建质量优于传统图像拼接与辐射场方法,对场景运动和非理想拍摄条件具有更强容忍度。

原文摘要 · Abstract (English)

Challenging to capture, and challenging to display on a cellphone screen, the panorama paradoxically remains both a staple and underused feature of modern mobile camera applications. In this work we address both of these challenges with a spherical neural light field model for implicit panoramic image stitching and re-rendering; able to accommodate for depth parallax, view-dependent lighting, and local scene motion and color changes during capture. Fit during test-time to an arbitrary path panoramic video capture -- vertical, horizontal, random-walk -- these neural light spheres jointly estimate the camera path and a high-resolution scene reconstruction to produce novel wide field-of-view projections of the environment. Our single-layer model avoids expensive volumetric sampling, and decomposes the scene into compact view-dependent ray offset and color components, with a total model size of 80 MB per scene, and real-time (50 FPS) rendering at 1080p resolution. We demonstrate improved reconstruction quality over traditional image stitching and radiance field methods, with significantly higher tolerance to scene motion and non-ideal capture settings.

全景拼接神经光场实时渲染移动视觉

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。