arXiv:2502.09278cs.CV2025-02中稿 · Pattern Recognitio…被引 3

用多视角扩散模型生成一致3D网格,解决图像转3D视角不统一问题。

ConsistentDreamer: View-Consistent Meshes Through Balanced Multi-View Gaussian Optimization

  • 先生成固定多视角图像,再用扩散模型采样中间视角,约束视图一致性。
  • 动态调整粗略形状与细节优化的权重,提升重建质量。
  • 适合需要高质量一致3D资产的虚拟仿真与生成任务。

最近的扩散模型进展显著提升了3D生成能力,使基于图像生成的资产可用于具身人工智能仿真。然而,图像到3D的多解性导致不同视角间内容与质量不一致,限制了其应用。现有方法通过从视图条件扩散先验采样视角优化3D模型,但扩散模型无法保证视图一致性。本文提出ConsistentDreamer:首先生成一组固定多视角先验图像,再利用另一扩散模型在它们之间采样随机视角,并通过得分蒸馏采样(SDS)损失约束视图差异,确保粗略形状一致。每轮迭代中,使用生成的多视角先验图像进行细节重建。为平衡粗略形状与细节优化,引入基于同方差不确定性的动态任务依赖权重,自动更新。同时采用透明度、深度畸变和法线对齐损失优化表面,便于网格提取。相比最先进方法,本方法在视图一致性和视觉质量上均有提升。

原文摘要 · Abstract (English)

Recent advances in diffusion models have significantly improved 3D generation, enabling the use of assets generated from an image for embodied AI simulations. However, the one-to-many nature of the image-to-3D problem limits their use due to inconsistent content and quality across views. Previous models optimize a 3D model by sampling views from a view-conditioned diffusion prior, but diffusion models cannot guarantee view consistency. Instead, we present ConsistentDreamer, where we first generate a set of fixed multi-view prior images and sample random views between them with another diffusion model through a score distillation sampling (SDS) loss. Thereby, we limit the discrepancies between the views guided by the SDS loss and ensure a consistent rough shape. In each iteration, we also use our generated multi-view prior images for fine-detail reconstruction. To balance between the rough shape and the fine-detail optimizations, we introduce dynamic task-dependent weights based on homoscedastic uncertainty, updated automatically in each iteration. Additionally, we employ opacity, depth distortion, and normal alignment losses to refine the surface for mesh extraction. Our method ensures better view consistency and visual quality compared to the state-of-the-art.

3D生成扩散模型多视角一致网格重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。