arXiv:2505.24245cs.CVcs.AI2025-05

融合扩散与自回归模型,实现多模态3D生成的高保真结构控制。

LTM3D: Bridging Token Spaces for Conditional 3D Generation with Auto-Regressive Diffusion Framework

  • 在令牌空间中结合扩散与自回归机制,提升形状依赖建模能力。
  • 通过重构引导采样降低不确定性,生成结构更精确的3D形状。
  • 支持多种3D表示(如SDF、点云、网格),适用于图文条件生成。

我们提出LTM3D,一种用于条件化3D形状生成的潜在令牌空间建模框架,融合了扩散模型与自回归(AR)模型的优势。尽管扩散方法擅长建模连续潜在空间,而自回归模型能有效捕捉令牌间依赖关系,但将两者结合用于3D生成仍具挑战。为此,LTM3D采用条件分布建模主干,利用掩码自编码器与扩散模型增强令牌依赖学习;引入前缀学习,使条件令牌与形状潜在令牌对齐,提升跨模态灵活性;并设计潜在令牌重建模块与重建引导采样,减少不确定性,提升生成形状的结构保真度。该方法运行于令牌空间,支持多种3D表示形式,包括符号距离场(SDF)、点云、网格和3D高斯泼溅(3D Gaussian Splatting)。在图像与文本条件下的3D生成任务中,实验表明LTM3D在提示保真度与结构准确性上均优于现有方法,且具备多模态、多表示的通用生成能力。

原文摘要 · Abstract (English)

We present LTM3D, a Latent Token space Modeling framework for conditional 3D shape generation that integrates the strengths of diffusion and auto-regressive (AR) models. While diffusion-based methods effectively model continuous latent spaces and AR models excel at capturing inter-token dependencies, combining these paradigms for 3D shape generation remains a challenge. To address this, LTM3D features a Conditional Distribution Modeling backbone, leveraging a masked autoencoder and a diffusion model to enhance token dependency learning. Additionally, we introduce Prefix Learning, which aligns condition tokens with shape latent tokens during generation, improving flexibility across modalities. We further propose a Latent Token Reconstruction module with Reconstruction-Guided Sampling to reduce uncertainty and enhance structural fidelity in generated shapes. Our approach operates in token space, enabling support for multiple 3D representations, including signed distance fields, point clouds, meshes, and 3D Gaussian Splatting. Extensive experiments on image- and text-conditioned shape generation tasks demonstrate that LTM3D outperforms existing methods in prompt fidelity and structural accuracy while offering a generalizable framework for multi-modal, multi-representation 3D generation.

3D生成扩散模型自回归令牌空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。