arXiv:2505.09140cs.CV2025-05被引 1

用拓扑信息提升3D点云生成质量,兼顾细节与结构一致性

TopoDiT-3D: Topology-Aware Diffusion Transformer with Bottleneck Structure for 3D Point Cloud Generation

  • 引入持久同调提取全局拓扑特征,通过瓶颈结构融合到生成过程
  • 在ShapeNet上实现比当前最佳模型更高的视觉质量和多样性,训练效率提升23%
  • 适合关注3D生成结构完整性的研究人员和工业应用开发者

近年来,扩散变压器(DiT)模型在3D点云生成方面取得显著进展。然而,现有方法主要关注局部特征提取,忽视了对形状一致性至关重要的全局拓扑信息(如空洞)。为此,我们提出TopoDiT-3D,一种具有瓶颈结构的拓扑感知扩散变压器。该模型利用Perceiver Resampler构建瓶颈结构,不仅实现了通过持久同调提取的拓扑信息与特征学习的融合,还能自适应过滤冗余局部特征,提升训练效率。实验表明,TopoDiT-3D在视觉质量、多样性及训练效率上均优于现有最先进模型。结果验证了丰富拓扑信息在3D点云生成中的关键作用及其与传统局部特征学习的协同效应。视频与代码已开源:https://github.com/Zechao-Guan/TopoDiT-3D。

原文摘要 · Abstract (English)

Recent advancements in Diffusion Transformer (DiT) models have significantly improved 3D point cloud generation. However, existing methods primarily focus on local feature extraction while overlooking global topological information, such as voids, which are crucial for maintaining shape consistency and capturing complex geometries. To address this limitation, we propose TopoDiT-3D, a Topology-Aware Diffusion Transformer with a bottleneck structure for 3D point cloud generation. Specifically, we design the bottleneck structure utilizing Perceiver Resampler, which not only offers a mode to integrate topological information extracted through persistent homology into feature learning, but also adaptively filters out redundant local features to improve training efficiency. Experimental results demonstrate that TopoDiT-3D outperforms state-of-the-art models in visual quality, diversity, and training efficiency. Furthermore, TopoDiT-3D demonstrates the importance of rich topological information for 3D point cloud generation and its synergy with conventional local feature learning. Videos and code are available at https://github.com/Zechao-Guan/TopoDiT-3D.

3D生成扩散模型拓扑感知点云

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。