arXiv:2507.07410cs.CV2025-07被引 1

一次完成物体缺失部分补全与新视角生成,速度快且效果好。

EscherNet++: Simultaneous Amodal Completion and Scalable View Synthesis through Masked Fine-Tuning and Enhanced Feed-Forward 3D Reconstruction

  • 用掩码微调让模型端到端完成补全与新视角生成。
  • 在遮挡任务上PSNR提升3.9,体积交并比提高0.28。
  • 适合需要快速3D重建的工业或真实场景应用。

我们提出EscherNet++,一种经过掩码微调的扩散模型,可零样本合成物体的新视角并具备非模态补全能力。现有方法需多阶段、复杂流程先补全缺失区域再生成新视角,忽视跨视角依赖,且存储与计算冗余。我们通过输入级和特征级掩码微调,实现端到端建模,显著提升新视角生成与补全性能。此外,无需额外训练即可与其它前馈图像到网格模型结合,在10输入设置下重建时间降低95%,达到领先效果。尽管微调时使用较小数据集和批次,模型仍实现最优性能,并成功推广至真实世界遮挡重建任务。

原文摘要 · Abstract (English)

We propose EscherNet++, a masked fine-tuned diffusion model that can synthesize novel views of objects in a zero-shot manner with amodal completion ability. Existing approaches utilize multiple stages and complex pipelines to first hallucinate missing parts of the image and then perform novel view synthesis, which fail to consider cross-view dependencies and require redundant storage and computing for separate stages. Instead, we apply masked fine-tuning including input-level and feature-level masking to enable an end-to-end model with the improved ability to synthesize novel views and conduct amodal completion. In addition, we empirically integrate our model with other feed-forward image-to-mesh models without extra training and achieve competitive results with reconstruction time decreased by 95%, thanks to its ability to synthesize arbitrary query views. Our method's scalable nature further enhances fast 3D reconstruction. Despite fine-tuning on a smaller dataset and batch size, our method achieves state-of-the-art results, improving PSNR by 3.9 and Volume IoU by 0.28 on occluded tasks in 10-input settings, while also generalizing to real-world occluded reconstruction.

3D重建图像补全扩散模型新视角合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。