用球面投影统一表示单图3D形状,解决视图不一致和结构表达难题。
SPGen: Spherical Projection as Consistent and Flexible Representation for Single Image 3D Shape Generation
- 将几何信息投影到包围球并展开为多层2D球面图,实现单一视角编码
- 在多个数据集上生成质量超越基线,且推理效率更高
- 适合需要高一致性与复杂结构生成的3D建模任务
现有单视图3D生成模型通常采用多视图扩散先验重建物体表面,但仍易出现视图间不一致,且难以准确表达复杂内部结构或非平凡拓扑。本文通过将几何信息投影至包围球并展开为紧凑、结构化的多层2D球面投影(SP)表示,实现了仅在图像域操作的生成。SPGen具备三大优势:(1) 一致性:单射球面映射以单一视角编码表面几何,天然消除视图不一致与歧义;(2) 灵活性:多层SP图可表示嵌套内部结构,并支持直接提升为封闭或开放的3D表面;(3) 效率:图像域形式使模型可直接继承强大2D扩散先验,支持低资源下的高效微调。大量实验表明,SPGen在几何质量与计算效率方面显著优于现有基线。
原文摘要 · Abstract (English)
Existing single-view 3D generative models typically adopt multiview diffusion priors to reconstruct object surfaces, yet they remain prone to inter-view inconsistencies and are unable to faithfully represent complex internal structure or nontrivial topologies. In particular, we encode geometry information by projecting it onto a bounding sphere and unwrapping it into a compact and structural multi-layer 2D Spherical Projection (SP) representation. Operating solely in the image domain, SPGen offers three key advantages simultaneously: (1) Consistency. The injective SP mapping encodes surface geometry with a single viewpoint which naturally eliminates view inconsistency and ambiguity; (2) Flexibility. Multi-layer SP maps represent nested internal structures and support direct lifting to watertight or open 3D surfaces; (3) Efficiency. The image-domain formulation allows the direct inheritance of powerful 2D diffusion priors and enables efficient finetuning with limited computational resources. Extensive experiments demonstrate that SPGen significantly outperforms existing baselines in geometric quality and computational efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。