arXiv:2512.25071cs.CV2025-12被引 2

仅用几张无姿态图片,一键完成高质量3D场景编辑。

Edit3r: Instant 3D Scene Editing from Sparse Unposed Images

  • 直接预测指令对齐的3D编辑,无需优化或姿态估计。
  • 在20个场景、100次编辑上实现更优语义对齐与3D一致性。
  • 适合需要实时3D编辑的用户,如游戏、虚拟现实开发者。

我们提出Edit3r,一种前馈式框架,可从无姿态、视角不一致的指令编辑图像中一次性重建并编辑3D场景。与需逐场景优化的方法不同,Edit3r直接预测与指令对齐的3D编辑,实现快速且逼真的渲染,无需优化或姿态估计。训练中的关键挑战在于缺乏多视角一致的编辑图像作为监督。为此,我们采用(i)基于SAM2的重着色策略生成跨视角一致的可靠监督信号,以及(ii)不对称输入策略,将重着色参考视图与原始辅助视图配对,促使网络融合并对齐差异观测。推理时,模型能有效处理由InstructPix2Pix等2D方法生成的编辑图像,即使训练中未见过此类编辑。为进行大规模定量评估,我们构建了DL3DV-Edit-Bench基准,基于DL3DV测试集,包含20个多样化场景、4类编辑类型,共100次编辑。综合定量与定性结果表明,Edit3r在语义对齐和3D一致性方面优于近期基线,且推理速度显著更快,适用于实时3D编辑应用。

原文摘要 · Abstract (English)

We present Edit3r, a feed-forward framework that reconstructs and edits 3D scenes in a single pass from unposed, view-inconsistent, instruction-edited images. Unlike prior methods requiring per-scene optimization, Edit3r directly predicts instruction-aligned 3D edits, enabling fast and photorealistic rendering without optimization or pose estimation. A key challenge in training such a model lies in the absence of multi-view consistent edited images for supervision. We address this with (i) a SAM2-based recoloring strategy that generates reliable, cross-view-consistent supervision, and (ii) an asymmetric input strategy that pairs a recolored reference view with raw auxiliary views, encouraging the network to fuse and align disparate observations. At inference, our model effectively handles images edited by 2D methods such as InstructPix2Pix, despite not being exposed to such edits during training. For large-scale quantitative evaluation, we introduce DL3DV-Edit-Bench, a benchmark built on the DL3DV test split, featuring 20 diverse scenes, 4 edit types and 100 edits in total. Comprehensive quantitative and qualitative results show that Edit3r achieves superior semantic alignment and enhanced 3D consistency compared to recent baselines, while operating at significantly higher inference speed, making it promising for real-time 3D editing applications.

3D编辑实时渲染图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。