按语义重要性排序3D形状令牌,让生成更高效准确
LoST: Level of Semantics Tokenization for 3D Shapes
- 根据语义显著性排序令牌,优先生成完整有意义的形状
- 重建精度超越现有方法,在几何与语义上均大幅领先
- 仅需1%-10%令牌即可完成高质量3D生成,适合下游任务
令牌化是多模态生成建模中的基础技术,尤其在自回归(AR)模型中至关重要,该类模型最近成为3D生成的有力选择。然而,3D形状的最佳令牌化方式仍是未解难题。当前最优方法主要依赖为渲染和压缩设计的几何层次细节(LoD)层级,这些空间层级通常令牌效率低且缺乏语义一致性,不利于自回归建模。我们提出语义层次令牌化(LoST),将令牌按语义显著性排序,使早期前缀能解码出具备主干语义的完整、合理形状,后续令牌则逐步细化实例相关的几何与语义细节。为训练LoST,我们引入关系间距离对齐(RIDA),一种新型3D语义对齐损失,将3D形状潜在空间的关系结构与语义DINO特征空间对齐。实验表明,LoST在重建性能上达到新标杆,显著优于基于LoD的现有3D形状令牌化方法,在几何与语义重建指标上均有大幅提升。此外,LoST实现了高效高质量的自回归3D生成,并支持语义检索等下游任务,仅需先前自回归模型所需令牌的0.1%-10%。
原文摘要 · Abstract (English)
Tokenization is a fundamental technique in the generative modeling of various modalities. In particular, it plays a critical role in autoregressive (AR) models, which have recently emerged as a compelling option for 3D generation. However, optimal tokenization of 3D shapes remains an open question. State-of-the-art (SOTA) methods primarily rely on geometric level-of-detail (LoD) hierarchies, originally designed for rendering and compression. These spatial hierarchies are often token-inefficient and lack semantic coherence for AR modeling. We propose Level-of-Semantics Tokenization (LoST), which orders tokens by semantic salience, such that early prefixes decode into complete, plausible shapes that possess principal semantics, while subsequent tokens refine instance-specific geometric and semantic details. To train LoST, we introduce Relational Inter-Distance Alignment (RIDA), a novel 3D semantic alignment loss that aligns the relational structure of the 3D shape latent space with that of the semantic DINO feature space. Experiments show that LoST achieves SOTA reconstruction, surpassing previous LoD-based 3D shape tokenizers by large margins on both geometric and semantic reconstruction metrics. Moreover, LoST achieves efficient, high-quality AR 3D generation and enables downstream tasks like semantic retrieval, while using only 0.1%-10% of the tokens needed by prior AR models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。