用扩散对齐稀疏表示,提升3D高斯点云生成质量
FLUX3D: High-Fidelity 3D Gaussian Generation with Diffusion-Aligned Sparse Representation

- 采用扩散对齐结构化潜变量,增强2D特征重建能力
- 在多个基准上显著超越现有方法,生成更清晰的3D高斯点云
- 适合需要高质量3D内容生成的研究者与开发者
稀疏体素表示已成为图像到3D高斯溅射(3DGS)生成的可扩展基础,但现有方法因两大结构性瓶颈难以保留输入图像的高频视觉细节。首先,其采用为语义抽象优化的判别性2D特征构建稀疏体素潜变量,抑制了重建线索并引入表示瓶颈;其次,在生成阶段,标准扩散变压器缺乏有效机制将密集2D图像标记与稀疏3D体素潜变量对齐,导致跨模态对应瓶颈。为此,我们提出FLUX3D,一个可扩展的图像到3DGS框架,通过改进表示学习与跨模态对齐来提升生成质量。我们重新审视基于稀疏体素的3D表示学习中的2D特征选择,提出扩散对齐结构化潜变量(DA-SLAT),并搭配解码器仅架构以提升3DGS重建保真度。同时设计稀疏结构感知扩散框架,集成稀疏结构多模态扩散变压器(SMDiT)与模态感知旋转位置编码(MARoPE),实现几何无关的2D-3D对齐。大量基准实验表明,FLUX3D在外观保真度上取得显著提升,大幅优于所有现有最先进(SOTA)方法。
原文摘要 · Abstract (English)
Sparse voxel representation has emerged as a scalable foundation for image-to-3D Gaussian Splatting (3DGS) generation, yet current methods struggle to preserve high-frequency visual details of input images due to two structural bottlenecks. First, they adopt discriminative 2D features optimized for semantic abstraction to construct sparse voxel latents, which suppress reconstructive cues and induce a representation bottleneck. Second, in the generation stage, standard diffusion transformers lack effective mechanisms to align dense 2D image tokens with sparse 3D voxel latents, resulting in a cross-modal correspondence bottleneck. To address these issues, we propose FLUX3D, a scalable image-to-3DGS framework that boosts both representation learning and cross-modal alignment during generation. We first revisit 2D feature selection for sparse-voxel-based 3D representation learning, propose Diffusion-Aligned Structured Latents (DA-SLAT) and couple it with a decoder-only architecture to improve 3DGS reconstruction fidelity. We also design a sparse-structure-aware diffusion framework, which integrates the Sparse-structure Multimodal Diffusion Transformer (SMDiT) and Modal-Aware Rotary Positional Embedding (MARoPE) to achieve geometry-agnostic 2D-3D alignment. Extensive benchmark experiments demonstrate that FLUX3D yields substantial improvements in appearance fidelity and significantly outperforms all state-of-the-art (SOTA) methods in generating high-quality 3DGS assets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。