用2D组合性思想优化3D高斯点,实现文本生成多物体场景。
CompGS: Unleashing 2D Compositionality for Compositional Text-to-3D via Dynamically Optimizing 3D Gaussians
- 将2D组合性迁移至3D高斯初始化,逐实体建模确保物体间合理交互。
- 动态调整空间参数,提升小物体细节生成能力,质量优于现有方法。
- 支持可控编辑与场景生成,适合3D内容创作与设计应用。
近期文本引导的图像生成突破推动了3D生成发展。尽管单个高质量3D物体生成已可行,但在3D空间中生成多个具有合理交互的物体(即组合式3D生成)仍面临重大挑战。本文提出CompGS,一种基于3D高斯泼溅(GS)的新型生成框架,实现高效组合式文本到3D内容生成。核心设计包括:(1) 基于2D组合性的3D高斯初始化:将成熟的2D组合性思想迁移至3D,逐实体初始化高斯参数,确保每个实体具有一致的3D先验并实现合理交互;(2) 动态优化策略:采用得分蒸馏采样(SDS)损失动态优化3D高斯。该方法可自动将3D高斯分解为独立实体部分,实现实体级与组合级联合优化,并通过动态调节各实体空间参数,适应不同尺度物体,显著提升小物体细节生成效果。在T3Bench上的定性与定量评估表明,CompGS在图像质量与语义对齐方面均优于现有方法。该框架还可轻松扩展至可控3D编辑,支持场景生成。
原文摘要 · Abstract (English)
Recent breakthroughs in text-guided image generation have significantly advanced the field of 3D generation. While generating a single high-quality 3D object is now feasible, generating multiple objects with reasonable interactions within a 3D space, a.k.a. compositional 3D generation, presents substantial challenges. This paper introduces CompGS, a novel generative framework that employs 3D Gaussian Splatting (GS) for efficient, compositional text-to-3D content generation. To achieve this goal, two core designs are proposed: (1) 3D Gaussians Initialization with 2D compositionality: We transfer the well-established 2D compositionality to initialize the Gaussian parameters on an entity-by-entity basis, ensuring both consistent 3D priors for each entity and reasonable interactions among multiple entities; (2) Dynamic Optimization: We propose a dynamic strategy to optimize 3D Gaussians using Score Distillation Sampling (SDS) loss. CompGS first automatically decomposes 3D Gaussians into distinct entity parts, enabling optimization at both the entity and composition levels. Additionally, CompGS optimizes across objects of varying scales by dynamically adjusting the spatial parameters of each entity, enhancing the generation of fine-grained details, particularly in smaller entities. Qualitative comparisons and quantitative evaluations on T3Bench demonstrate the effectiveness of CompGS in generating compositional 3D objects with superior image quality and semantic alignment over existing methods. CompGS can also be easily extended to controllable 3D editing, facilitating scene generation. We hope CompGS will provide new insights to the compositional 3D generation. Project page: https://chongjiange.github.io/compgs.html.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。