提出多尺度隐空间一致性模型,3D点云生成快100倍且质量更优。
Multi-scale Latent Point Consistency Models for 3D Shape Generation
- 分层隐空间表示从点到超点,融合多尺度信息增强去噪能力
- 生成速度提升100倍,同时在形状质量和多样性上超越现有模型
- 适合需要高效高质3D生成的工业设计与虚拟现实场景
一致性模型(CMs)显著加速了扩散模型的采样过程,在高分辨率图像合成中表现优异。为将这一进展拓展至基于点云的3D形状生成,我们提出一种新型多尺度隐空间点一致性模型(MLPCM)。MLPCM采用隐空间扩散框架,引入从点级到超点级的多层次隐表示,对应不同空间分辨率。设计多尺度隐空间整合模块与3D空间注意力机制,有效利用多级超点信息对点级隐表示进行条件去噪。此外,提出通过一致性蒸馏学习的隐空间一致性模型,将先验压缩为单步生成器,大幅提高采样效率,同时保持原教师模型性能。在标准基准ShapeNet和ShapeNet-Vol上的大量实验表明,MLPCM实现生成过程100倍加速,且在形状质量与多样性上均优于当前最先进扩散模型。
原文摘要 · Abstract (English)
Consistency Models (CMs) have significantly accelerated the sampling process in diffusion models, yielding impressive results in synthesizing high-resolution images. To explore and extend these advancements to point-cloud-based 3D shape generation, we propose a novel Multi-scale Latent Point Consistency Model (MLPCM). Our MLPCM follows a latent diffusion framework and introduces hierarchical levels of latent representations, ranging from point-level to super-point levels, each corresponding to a different spatial resolution. We design a multi-scale latent integration module along with 3D spatial attention to effectively denoise the point-level latent representations conditioned on those from multiple super-point levels. Additionally, we propose a latent consistency model, learned through consistency distillation, that compresses the prior into a one-step generator. This significantly improves sampling efficiency while preserving the performance of the original teacher model. Extensive experiments on standard benchmarks ShapeNet and ShapeNet-Vol demonstrate that MLPCM achieves a 100x speedup in the generation process, while surpassing state-of-the-art diffusion models in terms of both shape quality and diversity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。