用生成模型从单张线稿还原3D线框,误差仅5.3%。
Reconstruction of a 3D wireframe from a single line drawing via generative depth estimation

- 将3D重建转为条件深度估计,用潜在扩散模型处理投影模糊
- 在百万级图像-深度对上训练,平均深度误差5.3%
- 适合手绘线稿转3D建模,尤其对复杂形状表现稳健
将2D自由手绘草图转换为3D模型仍是计算机视觉中的关键挑战,有助于弥合流畅绘图与CAD之间的差距。传统单目深度重建方法不适用于线稿解析。我们提出一种生成式方法,将重建问题定义为条件密集深度估计任务。为此,采用带有条件框架的潜在扩散模型(LDM),以解决正交投影带来的固有歧义。模型在超过一百万对图像-深度数据上进行训练,跨不同形状复杂度表现出稳健性能,平均深度误差为5.3%。
原文摘要 · Abstract (English)
The conversion of 2D freehand sketches into 3D models remains a pivotal challenge in computer vision, bridging the gap between fluent sketching and CAD. Traditional monocular depth reconstruction techniques are not suitable for line drawing interpretation. We propose a generative approach by framing reconstruction as a conditional dense depth estimation task. To achieve this, we implemented a Latent Diffusion Model (LDM) with a conditioning framework to resolve the inherent ambiguities of orthographic projections. We trained our model using a dataset of over one million image-depth pairs. Our framework demonstrated robust performance across varying shape complexities, with 5.3 percent average depth error.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。