arXiv:2507.05499cs.CV2025-07

用共享潜空间生成16张一致的多视角图像,仅需15秒。

LoomNet: Enhancing Multi-View Image Generation via Latent Space Weaving

  • 多视角并行扩散,共享潜空间协同构建视图一致性。
  • 15秒生成16张高质多视角图像,优于现有方法。
  • 适合需要快速生成多样、一致3D视图的研究者。

从单张图像生成一致的多视角图像仍具挑战性,空间不一致性常导致表面重建中3D网格质量下降。为此,我们提出LoomNet,一种新型多视角扩散架构,通过多次并行应用同一扩散模型,协作构建并利用共享潜空间以实现视图一致性。每个视角推断生成对应新视角的编码,投影至三个正交平面;各平面内所有视角编码融合为单一聚合平面,再通过信息传播与缺失区域插值,整合为统一连贯的解释。最终潜空间用于渲染一致的多视角图像。LoomNet在15秒内生成16张高质量一致图像,在图像质量和重建指标上超越现有最优方法,且能从相同输入生成多样、合理的新增视角。

原文摘要 · Abstract (English)

Generating consistent multi-view images from a single image remains challenging. Lack of spatial consistency often degrades 3D mesh quality in surface reconstruction. To address this, we propose LoomNet, a novel multi-view diffusion architecture that produces coherent images by applying the same diffusion model multiple times in parallel to collaboratively build and leverage a shared latent space for view consistency. Each viewpoint-specific inference generates an encoding representing its own hypothesis of the novel view from a given camera pose, which is projected onto three orthogonal planes. For each plane, encodings from all views are fused into a single aggregated plane. These aggregated planes are then processed to propagate information and interpolate missing regions, combining the hypotheses into a unified, coherent interpretation. The final latent space is then used to render consistent multi-view images. LoomNet generates 16 high-quality and coherent views in just 15 seconds. In our experiments, LoomNet outperforms state-of-the-art methods on both image quality and reconstruction metrics, also showing creativity by producing diverse, plausible novel views from the same input.

多视角生成扩散模型潜空间3D重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。