用粗略布局生成可局部编辑的3D物体,无需文字提示。
CompoSE: Compositional Synthesis and Editing of 3D Shapes via Part-Aware Control

- 通过扩散变换器交替处理局部部件与全局上下文,实现分部控制。
- 在引导合成任务上显著优于现有方法,客观指标和大模型评估均领先。
- 支持部件替换、增删、保风格缩放等精细编辑,适合3D内容创作者。
生成与编辑高质量3D内容仍是计算机图形学的核心挑战。本文提出CompoSE,一种基于部件感知控制的3D形状组合合成与编辑方法。该方法输入一组代表不同部件的空间粗略几何体(如包围盒),输出具有分离部件的3D对象,支持对单个部件进行细粒度的组合式编辑。其核心思想是采用扩散变压器架构,交替执行局部部件处理与跨部件全局上下文聚合,并引入新颖的条件化技术确保强响应用户输入。重要的是,该方法能直接从用户的粗略布局引导中推断部件语义与对称性,无需部件级文本提示。大量实验表明,该方法在引导合成任务上显著优于现有方法,且在客观指标与大语言模型评估中表现优异。
原文摘要 · Abstract (English)
Creating and editing high-quality 3D content remains a central challenge in computer graphics. We address this challenge by introducing CompoSE, a novel method for Compositional Synthesis and Editing of 3D shapes via part-aware control. Our method takes as input a set of coarse geometric primitives (e.g., bounding boxes) that represent distinct object parts arranged in a particular spatial configuration, and synthesizes as output part-separated 3D objects that support localized granular (i.e., compositional) editing of individual parts. The key insight that enables our method is our use of a diffusion transformer architecture that alternates between processing each part locally and aggregating contextual information across parts globally, and features a novel conditioning technique that ensures strong adherence to the user's input. Importantly, our method learns to infer part semantics and symmetries directly from the user's coarse layout guidance, and does not require part-level text prompts. We demonstrate that our method enables powerful part-level editing capabilities, including context-aware substitution, addition, deletion, and style-preserving resizing operations. We show through extensive experiments that our method significantly outperforms existing approaches on guided synthesis, as measured by objective metrics and LLM-based evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。