arXiv:2509.00269cs.GRcs.CV2025-09被引 8

无需训练,直接在3D隐空间编辑模型,实现高保真精准修改。

3D-LATTE: Latent Space 3D Editing from Textual Instructions

  • 基于3D扩散模型隐空间操作,避免多视角不一致问题。
  • 融合3D注意力图与几何正则化,支持多种语义编辑任务。
  • 适合需要快速、精确3D内容修改的设计师和开发者使用。

尽管多视图扩散模型在文本/图像驱动的3D资产生成方面取得成功,但基于指令的3D资产编辑仍远落后于生成质量。主要原因是依赖2D先验的方法存在视角不一致的编辑信号问题。本文提出一种无需训练的编辑方法,直接在原生3D扩散模型的隐空间中操作,实现对3D几何结构的直接控制。通过融合生成过程中的3D注意力图与源对象信息,并结合几何感知正则化、傅里叶域谱调制策略以及3D增强优化步骤,本方法在多种形状和语义变换上均超越现有3D编辑技术,实现高保真、精准的编辑效果。

原文摘要 · Abstract (English)

Despite the recent success of multi-view diffusion models for text/image-based 3D asset generation, instruction-based editing of 3D assets lacks surprisingly far behind the quality of generation models. The main reason is that recent approaches using 2D priors suffer from view-inconsistent editing signals. Going beyond 2D prior distillation methods and multi-view editing strategies, we propose a training-free editing method that operates within the latent space of a native 3D diffusion model, allowing us to directly manipulate 3D geometry. We guide the edit synthesis by blending 3D attention maps from the generation with the source object. Coupled with geometry-aware regularization guidance, a spectral modulation strategy in the Fourier domain and a refinement step for 3D enhancement, our method outperforms previous 3D editing methods enabling high-fidelity and precise edits across a wide range of shapes and semantic manipulations. Our project webpage is https://mparelli.github.io/3d-latte

3D编辑扩散模型隐空间操作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。