0.4秒完成3D场景修复,实时交互更流畅。
InstaInpaint: Instant 3D-Scene Inpainting with Masked Large Reconstruction Model
- 用2D修复提议快速生成3D补全,无需复杂优化。
- 比之前方法快1000倍,且在两个基准上表现顶尖。
- 适合虚拟现实、物体插入等需要快速响应的应用。
近期3D场景重建技术已支持虚拟与增强现实中的实时观看。为提升沉浸感,需支持交互操作(如移动或编辑物体),因此提出3D场景补全方法以修复或补充修改后的几何结构。然而现有方法依赖耗时且计算密集的优化过程,难以用于实时或在线应用。本文提出InstaInpaint,一种基于参考的前馈式框架,可在0.4秒内从2D补全提议生成3D场景补全。我们设计自监督掩码微调策略,使自定义的大规模重建模型(LRM)在大规模数据集上成功训练。通过大量实验,分析并识别出若干关键设计,显著提升泛化能力、纹理一致性和几何正确性。InstaInpaint相较之前方法提速1000倍,同时在两个标准基准上保持最先进性能。此外,该方法在物体插入、多区域补全等灵活下游任务中也表现出良好泛化能力。更多视频结果见项目主页:https://dhmbb2.github.io/InstaInpaint_page/
原文摘要 · Abstract (English)
Recent advances in 3D scene reconstruction enable real-time viewing in virtual and augmented reality. To support interactive operations for better immersiveness, such as moving or editing objects, 3D scene inpainting methods are proposed to repair or complete the altered geometry. However, current approaches rely on lengthy and computationally intensive optimization, making them impractical for real-time or online applications. We propose InstaInpaint, a reference-based feed-forward framework that produces 3D-scene inpainting from a 2D inpainting proposal within 0.4 seconds. We develop a self-supervised masked-finetuning strategy to enable training of our custom large reconstruction model (LRM) on the large-scale dataset. Through extensive experiments, we analyze and identify several key designs that improve generalization, textural consistency, and geometric correctness. InstaInpaint achieves a 1000x speed-up from prior methods while maintaining a state-of-the-art performance across two standard benchmarks. Moreover, we show that InstaInpaint generalizes well to flexible downstream applications such as object insertion and multi-region inpainting. More video results are available at our project page: https://dhmbb2.github.io/InstaInpaint_page/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。