无需逐场景优化,实现多视角一致的3D编辑
FluSplat: Sparse-View 3D Editing without Test-Time Optimization

- 训练时用多视角几何约束正则化,保证跨视角一致性
- 推理仅需一次前向传播,速度比传统方法快数个数量级
- 适合需要快速、稳定3D编辑的交互式应用
文本引导图像编辑与3D高斯泼溅(3DGS)的进展使高质量3D场景操作成为可能。然而,现有方法依赖测试时迭代优化,在2D扩散编辑与3D重建间反复调整,计算成本高、场景依赖强,且易产生跨视角不一致。本文提出一种前馈式框架,实现稀疏视图下的跨视角一致3D编辑。不通过迭代3D精修来强制一致性,而是在训练阶段于图像域引入跨视角正则化。通过联合监督多视角编辑并施加几何对齐约束,模型在推理时无需每场景优化即可生成视图一致的结果。编辑后的视图通过前馈3DGS模型直接提升为3DGS表示,单次前向传播完成。实验表明,该方法在编辑保真度上具有竞争力,跨视角一致性显著优于基于优化的方法,且推理时间减少多个数量级。
原文摘要 · Abstract (English)
Recent advances in text-guided image editing and 3D Gaussian Splatting (3DGS) have enabled high-quality 3D scene manipulation. However, existing pipelines rely on iterative edit-and-fit optimization at test time, alternating between 2D diffusion editing and 3D reconstruction. This process is computationally expensive, scene-specific, and prone to cross-view inconsistencies. We propose a feed-forward framework for cross-view consistent 3D scene editing from sparse views. Instead of enforcing consistency through iterative 3D refinement, we introduce a cross-view regularization scheme in the image domain during training. By jointly supervising multi-view edits with geometric alignment constraints, our model produces view-consistent results without per-scene optimization at inference. The edited views are then lifted into 3D via a feedforward 3DGS model, yielding a coherent 3DGS representation in a single forward pass. Experiments demonstrate competitive editing fidelity and substantially improved cross-view consistency compared to optimization-based methods, while reducing inference time by orders of magnitude.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。