统一跨尺度3D生成与理解,速度提升21倍且精度大幅超越扩散模型。
Unified Cross-Scale 3D Generation and Understanding via Autoregressive Modeling
- 基于八叉树的分层压缩,将3D结构转为紧凑一维序列。
- 通过两级子树压缩,令牌数减少8倍,支持跨尺度建模。
- 适用于分子、蛋白质、晶体等多场景,推理快21.8倍,精度提升256%。
3D结构建模在不同尺度上至关重要,广泛应用于流体模拟、3D重建、蛋白质折叠和分子对接等领域。然而,现有方法仍高度专业化,缺乏跨任务和跨尺度的通用性。本文提出Uni-3DAR,一种统一的自回归框架,用于跨尺度3D生成与理解。核心是基于八叉树数据结构的粗到细分词器,将多样化的3D结构压缩为紧凑的一维令牌序列。进一步提出两级子树压缩策略,使八叉树令牌序列最多减少8倍。为应对压缩带来的动态令牌位置问题,引入掩码下一令牌预测机制,显著提升位置建模准确性,从而大幅增强模型性能。在小分子、蛋白质、聚合物、晶体及宏观3D物体等多个生成与理解任务上进行大量实验,验证其有效性与通用性。值得注意的是,Uni-3DAR在性能上显著超越先前最先进扩散模型,相对提升高达256%,同时推理速度最快达21.8倍提升。
原文摘要 · Abstract (English)
3D structure modeling is essential across scales, enabling applications from fluid simulation and 3D reconstruction to protein folding and molecular docking. Yet, despite shared 3D spatial patterns, current approaches remain fragmented, with models narrowly specialized for specific domains and unable to generalize across tasks or scales. We propose Uni-3DAR, a unified autoregressive framework for cross-scale 3D generation and understanding. At its core is a coarse-to-fine tokenizer based on octree data structures, which compresses diverse 3D structures into compact 1D token sequences. We further propose a two-level subtree compression strategy, which reduces the octree token sequence by up to 8x. To address the challenge of dynamically varying token positions introduced by compression, we introduce a masked next-token prediction strategy that ensures accurate positional modeling, significantly boosting model performance. Extensive experiments across multiple 3D generation and understanding tasks, including small molecules, proteins, polymers, crystals, and macroscopic 3D objects, validate its effectiveness and versatility. Notably, Uni-3DAR surpasses previous state-of-the-art diffusion models by a substantial margin, achieving up to 256\% relative improvement while delivering inference speeds up to 21.8x faster.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。