首个直接生成3D场景的框架,避免2D中间表示的失真问题。
Native3D: End-to-End 3D Scene Generation via Unified Mesh-Texture Modeling and Semantic Alignment

- 用统一网格纹理联合表示建模3D结构与纹理特征。
- 引入3D语义对齐损失,显著提升几何与纹理保真度。
- 适合需要高质量3D生成与灵活编辑的应用场景。
本文提出 Native3D,首个完全绕过2D中间表示的端到端3D场景生成框架。传统方法需将3D表示适配至2D域以利用预训练扩散模型,不可避免引入几何结构扭曲和纹理细节退化等域适应问题。为此,我们设计了一种基于Transformer的场景编码器,实现网格与纹理的联合建模,有效保持场景内物体间的空间关系与视觉一致性。进一步提出3D Representation Alignment Loss(3D REPA Loss),采用改进的对比学习机制对潜在空间中的多层级语义表示进行对齐,显著提升几何与纹理保真度。实验表明,Native3D在生成质量与编辑灵活性上均优于现有方法,为3D场景编辑提供了新解决方案。
原文摘要 · Abstract (English)
This paper presents Native3D, the first end-to-end 3D scene generation framework that completely bypasses 2D intermediate representations. Traditional approaches typically require adapting 3D representations to the 2D domain to leverage pre-trained diffusion models, which inevitably introduces domain adaptation issues including geometric structural distortion and texture detail degradation. To address these limitations, we design a unified mesh-texture joint representation that simultaneously models both geometric structures and texture features through a Transformer-based scene encoder, effectively maintaining spatial relationships and visual consistency among objects within scenes. We further propose the 3D Representation Alignment Loss (3D REPA Loss), which employs an improved contrastive learning mechanism to align multi-level semantic representations in the latent space, significantly enhancing geometric and textural fidelity. Experimental results demonstrate that Native3D outperforms existing methods in both generation quality and editing flexibility, providing a novel solution for 3D scene editing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。