arXiv:2602.05321cs.CV2026-02被引 1

直接处理广角镜头图像,实现360度全景三维重建

Wid3R: Wide Field-of-View 3D Reconstruction via Camera Model Conditioning

  • 用球谐函数建模光线,引入相机模型标记感知畸变
  • 在鱼眼和360数据集上AUC@30提升超77%
  • 首个支持多帧广角输入的前馈式3D重建方法

我们提出Wid3R,一种支持广角相机模型的多视角视觉几何重建前馈神经网络。与以往假设校正或针孔输入的方法不同,Wid3R可直接处理未校准、未去畸变的广角图像。该方法采用基于射线的表示,结合球谐函数,并引入新型相机模型标记以实现畸变感知重建。据我们所知,Wid3R是首个支持360°图像的多帧前馈3D重建方法。此外,通过在多种相机类型上进行条件化训练,显著提升了对360°场景的泛化能力,并缓解了数据稀疏问题。在Zip-NeRF(鱼眼)和Stanford2D3D(360)数据集上,其AUC@30分别提升达+33.67和+77.33。

原文摘要 · Abstract (English)

We present Wid3R, a feed-forward neural network for multi-view visual geometry reconstruction that supports wide field-of-view camera models. Unlike existing methods that assume rectified or pinhole inputs, Wid3R directly models wide-angle imagery without explicit calibration or undistortion. Our approach leverages a ray-based representation with spherical harmonics and introduces a novel camera model token to enable distortion-aware reconstruction. To the best of our knowledge, Wid3R is the first multi-frame feed-forward 3D reconstruction method that supports 360 imagery. Moreover, we show that conditioning on diverse camera types improves generalization to 360 scenes and alleviates data sparsity issues. Wid3R achieves significant performance gains, improving AUC@30 by up to +33.67 on Zip-NeRF (fisheye) and +77.33 on Stanford2D3D (360).

3D重建广角成像神经渲染相机建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。