直接在3D空间编辑,速度更快且保持形状一致性。
Native 3D Editing with Full Attention
- 用3D标记拼接代替交叉注意力,提升效率和精度。
- 在添加、删除、修改任务中表现优于现有2D升维方法。
- 适合需要快速高质量3D内容生成的设计师与开发者。
指令引导的3D编辑正快速发展,但现有方法存在显著局限:基于优化的方法过于缓慢,而依赖多视图2D编辑的前馈方法常导致几何不一致和视觉质量下降。为此,我们提出一种新型原生3D编辑框架,通过单次前馈过程直接操作3D表示。我们构建了一个大规模多模态数据集,涵盖多样化的添加、删除与修改任务,精心标注以确保编辑对象忠实响应指令,同时保持未编辑区域与源物体的一致性。基于该数据集,我们探索两种条件策略:传统交叉注意力机制与新颖的3D标记拼接方法。结果表明,标记拼接更参数高效且性能更优。大量实验显示,本方法在生成质量、3D一致性与指令保真度上均超越现有2D升维方法,树立新基准。
原文摘要 · Abstract (English)
Instruction-guided 3D editing is a rapidly emerging field with the potential to broaden access to 3D content creation. However, existing methods face critical limitations: optimization-based approaches are prohibitively slow, while feed-forward approaches relying on multi-view 2D editing often suffer from inconsistent geometry and degraded visual quality. To address these issues, we propose a novel native 3D editing framework that directly manipulates 3D representations in a single, efficient feed-forward pass. Specifically, we create a large-scale, multi-modal dataset for instruction-guided 3D editing, covering diverse addition, deletion, and modification tasks. This dataset is meticulously curated to ensure that edited objects faithfully adhere to the instructional changes while preserving the consistency of unedited regions with the source object. Building upon this dataset, we explore two distinct conditioning strategies for our model: a conventional cross-attention mechanism and a novel 3D token concatenation approach. Our results demonstrate that token concatenation is more parameter-efficient and achieves superior performance. Extensive evaluations show that our method outperforms existing 2D-lifting approaches, setting a new benchmark in generation quality, 3D consistency, and instruction fidelity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。