用统一模型同时生成图像颜色、深度和法向,效率更高更真实。
Orchid: Image Latent Diffusion for Joint Appearance and Geometry Generation
- 联合编码颜色、深度、法向到共享潜空间,一次扩散过程完成生成。
- 在法向预测和深度-法向一致性上超越现有专用模型。
- 支持文本生成、单目联合估计与大范围3D区域修复。
我们提出Orchid,一种统一的潜在扩散模型,通过学习联合外观-几何先验,在单一扩散过程中生成彩色图、深度图和表面法向图。该方法比当前分步使用独立模型的流程更高效且更一致。Orchid具有多功能性:可直接从文本生成彩色、深度和法向图像;支持以颜色为条件进行单目深度与法向联合估计微调;并能通过从联合分布采样,无缝修复大范围3D区域。其采用新型变分自编码器(VAE),将RGB、相对深度和表面法向联合编码至共享潜空间,并结合潜在扩散模型对这些潜变量进行去噪。大量实验表明,Orchid在几何预测任务上达到与顶尖专用方法相当的性能,甚至在法向预测准确率和深度-法向一致性方面表现更优。此外,它可联合修复彩色-深度-法向图像,生成效果比现有多步方法更具真实性。
原文摘要 · Abstract (English)
We introduce Orchid, a unified latent diffusion model that learns a joint appearance-geometry prior to generate color, depth, and surface normal images in a single diffusion process. This unified approach is more efficient and coherent than current pipelines that use separate models for appearance and geometry. Orchid is versatile - it directly generates color, depth, and normal images from text, supports joint monocular depth and normal estimation with color-conditioned finetuning, and seamlessly inpaints large 3D regions by sampling from the joint distribution. It leverages a novel Variational Autoencoder (VAE) that jointly encodes RGB, relative depth, and surface normals into a shared latent space, combined with a latent diffusion model that denoises these latents. Our extensive experiments demonstrate that Orchid delivers competitive performance against SOTA task-specific methods for geometry prediction, even surpassing them in normal-prediction accuracy and depth-normal consistency. It also inpaints color-depth-normal images jointly, with more qualitative realism than existing multi-step methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。