对比两种生成模型在单张深度图下高保真3D形状补全的表现。
Evaluating Latent Generative Paradigms for High-Fidelity 3D Shape Completion from a Single Depth Image
- 用扩散模型和自回归变换器完成3D形状补全,分别基于连续与离散潜在空间。
- 连续潜在扩散模型在噪声深度图下达到当前最优性能,优于判别式模型和自回归方法。
- 在相同离散潜空间下,自回归模型可媲美甚至超越扩散模型,适合特定场景。
尽管生成模型已在多种数据模态(包括3D数据)中广泛应用,但针对具体任务的最佳模型尚无共识。现有研究多依赖文本或图像等条件信息引导生成,而对部分3D数据作为条件的评估仍不充分。本文比较了两种最具前景的生成模型——去噪扩散概率模型(DDPM)和自回归因果变换器(Autoregressive Causal Transformers),并将其适配于生成式形状建模与补全任务。我们进行了全面的定量评估与对比,包含一个基线判别式模型及详尽的消融实验。结果表明:(1) 基于连续潜空间的扩散模型优于判别式模型与自回归方法,在真实条件下从单张噪声深度图进行多模态3D形状补全上达到当前最优性能;(2) 当在相同离散潜空间下比较时,自回归模型可在该任务上达到或超过扩散模型表现。
原文摘要 · Abstract (English)
While generative models have seen significant adoption across a wide range of data modalities, including 3D data, a consensus on which model is best suited for which task has yet to be reached. Further, conditional information such as text and images to steer the generation process are frequently employed, whereas others, like partial 3D data, have not been thoroughly evaluated. In this work, we compare two of the most promising generative models--Denoising Diffusion Probabilistic Models and Autoregressive Causal Transformers--which we adapt for the tasks of generative shape modeling and completion. We conduct a thorough quantitative evaluation and comparison of both tasks, including a baseline discriminative model and an extensive ablation study. Our results show that (1) the diffusion model with continuous latents outperforms both the discriminative model and the autoregressive approach and delivers state-of-the-art performance on multi-modal shape completion from a single, noisy depth image under realistic conditions and (2) when compared on the same discrete latent space, the autoregressive model can match or exceed diffusion performance on these tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。