用新型体素结构生成高质量3D资产,支持复杂拓扑与材质细节。
Native and Compact Structured Latents for 3D Generation

- 提出O-Voxel体素结构,统一编码几何与外观属性。
- 训练40亿参数流匹配模型,生成质量超越现有方法。
- 适合需要高精度3D生成的工业设计与影视制作场景。
近年来3D生成建模虽大幅提升生成真实感,但现有表示仍难以捕捉复杂拓扑与细节外观。本文提出一种从原生3D数据学习结构化隐空间的方法。核心是新型稀疏体素结构O-Voxel,一种能同时编码几何与外观的全向体素表示,可鲁棒建模开放、非流形及封闭表面,并捕获纹理颜色以外的物理渲染参数等综合表面属性。基于O-Voxel,设计稀疏压缩变分自编码器(Sparse Compression VAE),实现高空间压缩率与紧凑隐空间。使用多样公共3D资产数据集训练大规模40亿参数流匹配模型用于3D生成,推理高效,生成资产的几何与材质质量显著优于现有模型。本方法为3D生成建模带来重要进展。
原文摘要 · Abstract (English)
Recent advancements in 3D generative modeling have significantly improved the generation realism, yet the field is still hampered by existing representations, which struggle to capture assets with complex topologies and detailed appearance. This paper present an approach for learning a structured latent representation from native 3D data to address this challenge. At its core is a new sparse voxel structure called O-Voxel, an omni-voxel representation that encodes both geometry and appearance. O-Voxel can robustly model arbitrary topology, including open, non-manifold, and fully-enclosed surfaces, while capturing comprehensive surface attributes beyond texture color, such as physically-based rendering parameters. Based on O-Voxel, we design a Sparse Compression VAE which provides a high spatial compression rate and a compact latent space. We train large-scale flow-matching models comprising 4B parameters for 3D generation using diverse public 3D asset datasets. Despite their scale, inference remains highly efficient. Meanwhile, the geometry and material quality of our generated assets far exceed those of existing models. We believe our approach offers a significant advancement in 3D generative modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。