首次融合自回归与扩散模型,实现3D理解与生成统一
PnP-U3D: Plug-and-Play 3D Framework Bridging Autoregression and Diffusion for Unified Understanding and Generation
- 用自回归做3D理解,扩散模型做3D生成,中间用轻量Transformer桥接
- 在多个3D任务上达到当前最佳性能,且支持高效编辑
- 适合研究3D通用智能或跨模态生成的学者
大型多模态模型的快速发展推动了理解与生成统一框架的研究。尽管2D领域已取得显著进展,3D方向仍处于探索阶段。现有基于自回归范式的统一方法因强制信号量化和高昂训练成本导致性能大幅下降。本文核心洞察在于:关键不在于强求统一自回归范式,而在于实现生成与理解间高效的信息交互,同时最小化对各自能力的损害,并利用预训练模型降低训练开销。为此,我们提出首个融合自回归与扩散的3D统一框架:采用自回归进行3D理解,连续扩散进行3D生成;通过轻量Transformer连接大语言模型特征空间与3D扩散模型条件空间,实现跨模态有效信息交换,同时保留独立模型所学先验。大量实验表明,该框架在多种3D理解与生成基准上达到领先水平,并在3D编辑任务中表现优异。结果表明,统一的AR+扩散模型是构建通用3D智能的有前途方向。
原文摘要 · Abstract (English)
The rapid progress of large multimodal models has inspired efforts toward unified frameworks that couple understanding and generation. While such paradigms have shown remarkable success in 2D, extending them to 3D remains largely underexplored. Existing attempts to unify 3D tasks under a single autoregressive (AR) paradigm lead to significant performance degradation due to forced signal quantization and prohibitive training cost. Our key insight is that the essential challenge lies not in enforcing a unified autoregressive paradigm, but in enabling effective information interaction between generation and understanding while minimally compromising their inherent capabilities and leveraging pretrained models to reduce training cost. Guided by this perspective, we present the first unified framework for 3D understanding and generation that combines autoregression with diffusion. Specifically, we adopt an autoregressive next-token prediction paradigm for 3D understanding, and a continuous diffusion paradigm for 3D generation. A lightweight transformer bridges the feature space of large language models and the conditional space of 3D diffusion models, enabling effective cross-modal information exchange while preserving the priors learned by standalone models. Extensive experiments demonstrate that our framework achieves state-of-the-art performance across diverse 3D understanding and generation benchmarks, while also excelling in 3D editing tasks. These results highlight the potential of unified AR+diffusion models as a promising direction for building more general-purpose 3D intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。