arXiv:2606.25430cs.CV2026-06

无需迭代采样,36秒完成单图3D重建。

PRISM: Feed-Forward Single-Image 3D Reconstruction via Geometric Warp-Residual Modeling

论文配图:PRISM: Feed-Forward Single-Image 3D Reconstruction via Geometric Warp-Residual Modeling
图 1 · 摘自论文原文
  • 用几何形变预估主结构,仅对残差学习修正。
  • 在三个基准上达到扩散模型水平的重建质量。
  • 适合需要快速生成3D内容的场景应用。

从单张图像重建3D场景是计算机视觉的基础挑战,广泛应用于虚拟现实、机器人和内容创作。现有方法借助相机控制的视频扩散模型取得优异性能,但依赖迭代扩散采样,严重限制实际部署。我们观察到,仅靠几何前向形变即可直接覆盖目标视图的大部分内容,仅剩少量紧凑残差需编码器修正。受此启发,提出PRISM,一种前馈框架,将多视角潜在预测分解为无参数几何先验与学习残差修正,推理无需扩散采样。为实现纯合成数据训练下的泛化能力,设计两阶段训练策略:先通过潜变量监督蒸馏提升几何泛化性,再通过感知微调优化外观质量。在三个基准上的实验表明,PRISM在重建质量上可媲美扩散方法,同时推理时间大幅降低至每场景仅36秒。

原文摘要 · Abstract (English)

Reconstructing 3D scenes from a single image is a fundamental challenge in computer vision, with broad applications in virtual reality, robotics, and content creation. Recent methods achieve outstanding performance by leveraging camera-controlled video diffusion models, but rely on iterative diffusion sampling, which greatly limits their practical deployment. We observe that geometric forward warping alone can cover the majority of a target view directly from the input image, with only a compact residual left for the encoder to correct. Motivated by this observation, we propose PRISM, a feed-forward framework that decomposes multi-view latent prediction into a parameter-free geometric prior and a learned residual correction, with no diffusion sampling required at inference. To enable generalization from purely synthetic training data, we devise a two-stage training strategy combining latents supervised distillation for geometric generalization and perceptual fine-tuning for appearance quality optimization. Extensive experiments on three benchmarks demonstrate that PRISM achieves competitive reconstruction quality compared with diffusion-based methods, while reducing inference time dramatically to only 36 seconds per scene.

3D重建单图建模前馈架构几何形变

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。