让AI根据文字描述一次生成多个3D物体,位置和姿态更可控。
COMOGen: A Controllable Text-to-3D Multi-object Generation Framework
- 通过布局与多视角先验知识蒸馏实现多物体协同生成。
- 在ShapeNet等数据集上生成质量优于当前最优方法。
- 适合需要精确控制多个3D物体位置的场景设计者使用。
现有文本到3D物体生成方法主要针对单个物体描述,难以准确控制多物体场景中的位置关系。为此,本文提出COMOGen——一种可控制的文本到多物体3D生成框架。该框架通过布局控制模块、多视角一致性控制模块和3D内容增强模块,融合布局与多视角先验知识,实现多个3D物体的同时生成。为统一两类先验知识,提出布局-多视角评分蒸馏(Layout Multi-view Score Distillation),进一步提升生成内容的多样性与质量。大量实验表明,该方法在多个基准数据集上均显著优于当前最先进方法,推动了可控文本驱动3D内容生成的发展。
原文摘要 · Abstract (English)
The controllability of 3D object generation methods is achieved through input text. Existing text-to-3D object generation methods primarily focus on generating a single object based on a single object description. However, these methods often face challenges in producing results that accurately correspond to our desired positions when the input text involves multiple objects. To address the issue of controllability in generating multiple objects, this paper introduces COMOGen, a COntrollable text-to-3D Multi-Object Generation framework. COMOGen enables the simultaneous generation of multiple 3D objects by the distillation of layout and multi-view prior knowledge. The framework consists of three modules: the layout control module, the multi-view consistency control module, and the 3D content enhancement module. Moreover, to integrate these three modules as an integral framework, we propose Layout Multi-view Score Distillation, which unifies two prior knowledge and further enhances the diversity and quality of generated 3D content. Comprehensive experiments demonstrate the effectiveness of our approach compared to the state-of-the-art methods, which represents a significant step forward in enabling more controlled and versatile text-based 3D content generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。