用神经场模型直接从手机拍摄数据重建三维场景,无需标注或复杂预处理。
Neural Field Representations of Mobile Computational Photography
- 设计神经场模型,将手机传感器采集的原始数据映射为连续空间表示
- 在真实场景数据上实现深度估计与图像拼接,性能超越现有方法
- 适合移动计算摄影、无监督3D重建方向的研究者参考
过去二十年间,智能手机影像技术飞速发展,已取代其他数字摄影形式。现代手机集成激光测距、多焦距摄像头阵列、分像素传感器及陀螺仪、加速度计、磁力计等非视觉传感器,并配备专用芯片进行图像与信号处理,成为便携式计算成像平台。近年来,神经场——即通过训练小规模神经网络将连续空间坐标映射到输出信号——被用于无需显式数据表示(如像素阵列或点云)即可重建复杂场景。本论文展示,经精心设计的神经场模型可紧凑表达复杂几何与光照效果,直接从野外手机拍摄数据中实现深度估计、层分离与图像拼接。这些方法在不依赖复杂预处理、标签真值数据或机器学习先验的情况下,表现优于当前最优方案。其核心是利用结构良好的自正则化模型,通过随机梯度下降拟合智能手机原始测量数据,直接求解困难的逆问题。
原文摘要 · Abstract (English)
Over the past two decades, mobile imaging has experienced a profound transformation, with cell phones rapidly eclipsing all other forms of digital photography in popularity. Today's cell phones are equipped with a diverse range of imaging technologies - laser depth ranging, multi-focal camera arrays, and split-pixel sensors - alongside non-visual sensors such as gyroscopes, accelerometers, and magnetometers. This, combined with on-board integrated chips for image and signal processing, makes the cell phone a versatile pocket-sized computational imaging platform. Parallel to this, we have seen in recent years how neural fields - small neural networks trained to map continuous spatial input coordinates to output signals - enable the reconstruction of complex scenes without explicit data representations such as pixel arrays or point clouds. In this thesis, I demonstrate how carefully designed neural field models can compactly represent complex geometry and lighting effects. Enabling applications such as depth estimation, layer separation, and image stitching directly from collected in-the-wild mobile photography data. These methods outperform state-of-the-art approaches without relying on complex pre-processing steps, labeled ground truth data, or machine learning priors. Instead, they leverage well-constructed, self-regularized models that tackle challenging inverse problems through stochastic gradient descent, fitting directly to raw measurements from a smartphone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。