arXiv:2503.08382cs.CV2025-03CVPR被引 2

仅用几张照片就能还原场景光照与物体材质,速度快且精度高。

Twinner: Shining Light on Digital Twins in a Few Snaps

论文配图:Twinner: Shining Light on Digital Twins in a Few Snaps
图 1 · 摘自论文原文
  • 用高效体素变换器实现低内存占用的三维重建
  • 在斯坦福ORB数据集上效果超越现有网络,接近慢速优化方法
  • 基于合成数据训练+可微物理着色微调,适合真实场景应用

我们提出首个大规模重建模型 Twinner,仅需少量带姿态图像即可恢复场景光照、物体几何与材质属性。该模型基于大尺度重建框架,在三个方面创新:1)引入内存效率高的体素网格变换器,内存消耗随体素网格规模二次增长;2)构建大规模全合成数据集,包含程序生成的PBR纹理物体及多样的光照条件;3)通过可微物理着色模型在真实数据上微调,无需真实光照或材质标签,有效缩小了合成到真实的差距。我们在真实场景的斯坦福ORB基准测试中验证了模型有效性:仅用少量输入视图,重建质量显著优于现有前馈网络,接近更慢的逐场景优化方法。

原文摘要 · Abstract (English)

We present the first large reconstruction model, Twinner, capable of recovering a scene's illumination as well as an object's geometry and material properties from only a few posed images. Twinner is based on the Large Reconstruction Model and innovates in three key ways: 1) We introduce a memory-efficient voxel-grid transformer whose memory scales only quadratically with the size of the voxel grid. 2) To deal with scarcity of high-quality ground-truth PBR-shaded models, we introduce a large fully-synthetic dataset of procedurally-generated PBR-textured objects lit with varied illumination. 3) To narrow the synthetic-to-real gap, we finetune the model on real life datasets by means of a differentiable physically-based shading model, eschewing the need for ground-truth illumination or material properties which are challenging to obtain in real life. We demonstrate the efficacy of our model on the real life StanfordORB benchmark where, given few input views, we achieve reconstruction quality significantly superior to existing feedforward reconstruction networks, and comparable to significantly slower per-scene optimization methods.

三维重建数字孪生生成模型物理渲染

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。