LATTICE实现高保真3D资产的规模化生成,突破质量和效率瓶颈。
LATTICE: Democratize High-Fidelity 3D Generation at Scale
- 提出VoxSet半结构化表示,用粗粒度体素网格压缩3D资产为紧凑潜在向量
- 两阶段生成:先稀疏体素化锚点,再用修正流Transformer细化几何结构
- 支持任意分辨率解码、低成本训练,适合大规模3D内容创作场景
我们提出LATTICE,一种高保真3D资产生成框架,弥合了2D与3D生成模型在质量和可扩展性之间的差距。尽管2D图像生成得益于固定的网格结构和成熟的Transformer架构,3D生成仍因需从零预测空间结构与精细几何表面而更具挑战性,且现有3D表示计算复杂,缺乏结构化可扩展的编码方案。为此,我们提出VoxSet,一种半结构化表示,将3D资产压缩为锚定在粗粒度体素网格上的紧凑潜在向量集,实现高效且位置感知的生成。VoxSet保留了早期VecSet方法的简洁性与压缩优势,并在潜在空间中引入显式结构,使位置嵌入能引导生成并支持强大的令牌级推理时缩放。基于此表示,LATTICE采用两阶段流程:首先生成稀疏体素化几何锚点,随后利用修正流Transformer生成详细几何结构。该方法核心简单,支持任意分辨率解码、低成本训练和灵活推理,多项指标达到当前最优,在实现可扩展高保真3D内容生成方面迈出重要一步。
原文摘要 · Abstract (English)
We present LATTICE, a new framework for high-fidelity 3D asset generation that bridges the quality and scalability gap between 3D and 2D generative models. While 2D image synthesis benefits from fixed spatial grids and well-established transformer architectures, 3D generation remains fundamentally more challenging due to the need to predict both spatial structure and detailed geometric surfaces from scratch. These challenges are exacerbated by the computational complexity of existing 3D representations and the lack of structured and scalable 3D asset encoding schemes. To address this, we propose VoxSet, a semi-structured representation that compresses 3D assets into a compact set of latent vectors anchored to a coarse voxel grid, enabling efficient and position-aware generation. VoxSet retains the simplicity and compression advantages of prior VecSet methods while introducing explicit structure into the latent space, allowing positional embeddings to guide generation and enabling strong token-level test-time scaling. Built upon this representation, LATTICE adopts a two-stage pipeline: first generating a sparse voxelized geometry anchor, then producing detailed geometry using a rectified flow transformer. Our method is simple at its core, but supports arbitrary resolution decoding, low-cost training, and flexible inference schemes, achieving state-of-the-art performance on various aspects, and offering a significant step toward scalable, high-quality 3D asset creation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。