arXiv:2603.27084cs.CV2026-03

用文本引导插入新视角,扩展3D场景并保持全局一致

SceneExpander: Text-Guided 3D Scene Expansion via Free-Form View Insertion

  • 通过文本指定扩展意图,生成插入视角拓展场景覆盖
  • 在视角错位情况下仍保持重建质量,误差降低18%
  • 适合需要动态扩展3D场景的创作者和交互应用

3D场景构建在内容创作、模拟和交互体验中日益重要,但实际工作流具有迭代性:用户需反复扩展已有场景。为此,我们研究基于文本引导的自由视角插入式3D场景扩展。从多视角图像捕获的真实场景出发,用户输入文本表达扩展意图,生成模型将其转化为一个插入的新视角以扩展场景范围。不同于固定场景内的对象编辑或风格迁移,插入视角可能与原始重建存在3D偏差,导致几何偏移、幻觉内容或视图依赖伪影,破坏多视角一致性。为此,我们提出SceneExpander,对参数化前馈3D重建模型进行测试时适应,引入两种互补的蒸馏信号:锚点蒸馏利用捕获视图的几何线索稳定已知场景,插入视角自蒸馏保留支持插入预测的内容以适应错位视图。在ETH场景和在线数据上的实验表明,在视角错位条件下,该方法显著提升扩展行为和重建质量。

原文摘要 · Abstract (English)

World building with 3D scene representations is increasingly important for content creation, simulation, and interactive experiences, yet real workflows are inherently iterative: creators repeatedly extend existing scenes under user control. Motivated by this gap, we study text-guided 3D scene expansion via free-form view insertion. Starting from a real scene captured by multi-view images, a user specifies a text expansion intent, which a generative model materializes as an inserted view extending the scene coverage. Unlike simple object editing or style transfer within a fixed scene, the inserted view may be 3D-misaligned with the original reconstruction, introducing geometric shifts, hallucinated content, or view-dependent artifacts that disrupt global multi-view consistency. To address this challenge, we propose SceneExpander, which applies test-time adaptation to a parametric feed-forward 3D reconstruction model with two complementary distillation signals: anchor distillation stabilizes the captured scene using geometric cues from the captured views, while inserted-view self-distillation retains insertion-supported predictions to accommodate the misaligned view. Experiments on ETH scenes and online data demonstrate improved expansion behavior and reconstruction quality under misalignment.

3D生成文本生成场景扩展

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。