解决第一人称3D场景生成中的视角不一致与几何失真问题
CGGS: Consistency-Augmented Geometric Gaussian Splatting for Ego-Centric 3D Scene Generation

- 通过一致性增强损失微调多视角扩散模型生成一致高保真2D图像
- 利用光流与点追踪估计深度,从2D先验构建稠密点云粗布局
- 引入基于互信息的深度损失和分层优化,提升3D高斯重建质量
由于视角重叠有限且个体视角主导场景理解,第一人称3D场景生成仍面临挑战。本文提出CGGS,一种文本到3D框架,旨在增强3D内容感知并缓解几何失真。首先,通过一致性增强损失微调多视图潜在扩散模型,构建第一人称生成器,生成与文本描述对齐的一致、高保真2D内容。随后,布局装饰器利用光流与点轨迹对应关系估计深度,从第一人称2D先验生成稠密点云作为粗略布局。在此基础上,几何精修器采用基于熵的互信息深度损失(MID)结合分层优化方案,提升3D高斯重建的视觉质量和几何结构准确性。大量实验表明,CGGS在生成连贯、准确的文本驱动3D场景方面优于现有方法。
原文摘要 · Abstract (English)
Challenges remain in ego-centric 3D scene generation due to limited view overlap and the dominant influence of individual perspectives on scene interpretation. These factors hinder the creation of viewpoint-consistent and semantically aligned visual content, as well as the construction of accurate geometric structures. In this paper, we propose CGGS, a text-to-3D framework aiming to enhance 3D-content-awareness and address geometric distortions in ego-centric scene generation. Firstly, the Ego-centric Generator is proposed by fine-tuning a Multi-View Latent Diffusion Model with consistency-augmented loss to generate consistent, high-fidelity 2D content aligned with textual descriptions. Then, Layout Decorator leverages optical flow and point-track correspondence to estimate depth, therefore producing dense point clouds as coarse layouts from the ego-centric 2D priors. Building on this initialization, Geometric Refiner is proposed to enhance 3D Gaussian reconstruction via an entropy-based Mutual Information Depth Loss (MID) combined with a hierarchical optimization scheme for improving visual quality and geometric structure. Comprehensive experiments demonstrate that CGGS outperforms previous methods in generating coherent and accurate text-driven 3D scenes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。