从二维图生成立体3D模型,提升体积感与细节质量
REVIVE 3D: Refinement via Encoded Voluminous Inflated prior for Volume Enhancement

- 先膨胀轮廓补全整体体积,再注入局部细节
- 通过噪声反演增强几何结构,显著提升体积感
- 支持图像编辑,适合需要真实感3D建模的场景
近期生成模型在从2D图像生成多样化3D资产方面表现优异,但面对提供有限3D线索的平面图像时,仍难以生成具有充足体积感的3D内容。本文提出REVIVE 3D,一个两阶段、即插即用的3D生成流水线。第一阶段通过膨胀前景轮廓恢复全局体积,并叠加部件感知细节以捕捉局部结构;第二阶段采用3D潜空间精炼,向膨胀先验的潜空间注入高斯噪声并逐步去噪,利用几何线索激发骨干模型预训练的3D知识。此外,该框架支持图像条件下的3D编辑。为量化体积饱满度与表面平坦性,提出紧凑度(Compactness)和法向各向异性(Normal Anisotropy)指标。用户研究验证了这些指标与人类对体积感和质量的感知高度一致。在挑战性平面图像数据集上,基于大量定性和定量评估,REVIVE 3D达到当前最优性能。
原文摘要 · Abstract (English)
Recent generative models have shown strong performance in generating diverse 3D assets from 2D images, a fundamental research topic in computer vision and graphics. However, these models still struggle to generate voluminous 3D assets when the input is a flat image that provides limited 3D cues. We introduce REVIVE 3D, a two-stage, plug-and-play pipeline for generating voluminous 3D assets from flat images. In Stage 1, we construct an Inflated Prior by inflating the foreground silhouette to recover global volume and superimposing part-aware details to capture local structure. In Stage 2, 3D Latent Refinement injects Gaussian noise into the Inflated Prior's latent and then denoises it, using the prior's geometric cues to leverage the backbone's pretrained 3D knowledge. Furthermore, our framework supports image-conditioned 3D editing. To quantify volume and surface flatness, we propose Compactness and Normal Anisotropy. We validate Compactness and Normal Anisotropy through a user study, showing that these metrics align with human perception of volume and quality. We show that REVIVE 3D achieves state-of-the-art performance on a challenging flat image dataset, based on extensive qualitative and quantitative evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。