用多个解耦特征场实现3D内容的语义与结构化编辑
Structurally Disentangled Feature Fields Distillation for 3D Understanding and Editing
- 将3D特征拆分为视图相关与无关的多个解耦场
- 仅需2D监督即可学习,支持细粒度控制与编辑
- 适合需要精准3D内容操作的研究者与开发者
近期工作证明,可将使用大规模预训练2D模型获得的2D特征迁移到3D,仅依赖2D监督即可实现强大的3D理解与编辑能力。然而,现有方法通常假设3D特征由单一特征场捕获,并简化为视图无关。本文提出使用多个解耦特征场来捕捉3D特征的不同结构成分,包括视图相关与视图无关部分,这些均可仅通过2D特征监督学习得到。每个成分可独立控制,从而实现语义与结构化的理解与编辑。例如,用户点击后可分割出特定物体的3D特征,并独立编辑其视图相关的(反射)属性。我们在3D分割任务上评估该方法,展示了若干新颖的理解与编辑任务。
原文摘要 · Abstract (English)
Recent work has demonstrated the ability to leverage or distill pre-trained 2D features obtained using large pre-trained 2D models into 3D features, enabling impressive 3D editing and understanding capabilities using only 2D supervision. Although impressive, models assume that 3D features are captured using a single feature field and often make a simplifying assumption that features are view-independent. In this work, we propose instead to capture 3D features using multiple disentangled feature fields that capture different structural components of 3D features involving view-dependent and view-independent components, which can be learned from 2D feature supervision only. Subsequently, each element can be controlled in isolation, enabling semantic and structural understanding and editing capabilities. For instance, using a user click, one can segment 3D features corresponding to a given object and then segment, edit, or remove their view-dependent (reflective) properties. We evaluate our approach on the task of 3D segmentation and demonstrate a set of novel understanding and editing tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。