arXiv:2506.23205cs.CV2025-06AAAI被引 2

用最优传输理论提升3D形状补全的完整性与细节质量

BridgeShape: Latent Diffusion Schrödinger Bridge for 3D Shape Completion

  • 将补全过程建模为最优传输路径,确保全局结构一致性
  • 在压缩潜空间中操作,支持高分辨率生成且保持几何精度
  • 适合需要精细3D重建的工业设计与逆向工程场景

现有基于扩散的3D形状补全方法多采用条件范式,通过深层特征交互(如拼接、交叉注意力)将不完整形状信息注入去噪网络,引导采样生成完整形状,通常以体素距离函数表示。然而,这些方法未能显式建模最优全局传输路径,导致补全效果欠佳。此外,直接在体素空间进行扩散会带来分辨率限制,难以生成精细几何细节。为此,本文提出BridgeShape,一种基于潜空间扩散薛定谔桥的3D形状补全新框架。核心创新在于:(i) 将形状补全视为最优传输问题,显式建模不完整与完整形状间的转换路径,确保全局一致变换;(ii) 引入深度增强型向量量化变分自编码器(VQ-VAE),利用自投影多视角深度信息结合强DINOv2特征,增强几何结构感知能力。在紧凑但结构信息丰富的潜空间中操作,有效缓解分辨率限制,实现更高效、高保真度的3D形状补全。BridgeShape在大规模3D形状补全基准上达到当前最优性能,尤其在高分辨率及未见物体类别上表现优异。

原文摘要 · Abstract (English)

Existing diffusion-based 3D shape completion methods typically use a conditional paradigm, injecting incomplete shape information into the denoising network via deep feature interactions (e.g., concatenation, cross-attention) to guide sampling toward complete shapes, often represented by voxel-based distance functions. However, these approaches fail to explicitly model the optimal global transport path, leading to suboptimal completions. Moreover, performing diffusion directly in voxel space imposes resolution constraints, limiting the generation of fine-grained geometric details. To address these challenges, we propose BridgeShape, a novel framework for 3D shape completion via latent diffusion Schrödinger bridge. The key innovations lie in two aspects: (i) BridgeShape formulates shape completion as an optimal transport problem, explicitly modeling the transition between incomplete and complete shapes to ensure a globally coherent transformation. (ii) We introduce a Depth-Enhanced Vector Quantized Variational Autoencoder (VQ-VAE) to encode 3D shapes into a compact latent space, leveraging self-projected multi-view depth information enriched with strong DINOv2 features to enhance geometric structural perception. By operating in a compact yet structurally informative latent space, BridgeShape effectively mitigates resolution constraints and enables more efficient and high-fidelity 3D shape completion. BridgeShape achieves state-of-the-art performance on large-scale 3D shape completion benchmarks, demonstrating superior fidelity at higher resolutions and for unseen object classes.

3D补全扩散模型潜空间最优传输

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。