arXiv:2602.06158cs.CV2026-02被引 1

融合视觉与几何先验,提升复杂场景单视角3D重建精度

MGP-KAD: Multimodal Geometric Priors and Kolmogorov-Arnold Decoder for Single-View 3D Reconstruction in Complex Scenes

论文配图:MGP-KAD: Multimodal Geometric Priors and Kolmogorov-Arnold Decoder for Single-View 3D Reconstruction in Complex Scenes
图 1 · 摘自论文原文
  • 引入类级几何先验,动态调整以增强形状理解
  • 采用KAN混合解码器,更好处理多模态输入的复杂关系
  • 在Pix3D上实现当前最优性能,尤其改善细节与光滑性

复杂真实场景中的单视角3D重建面临噪声、物体多样性及数据集有限等挑战。为此,我们提出MGP-KAD框架,通过融合RGB图像与几何先验信息,提升重建准确性。几何先验基于真实物体数据采样与聚类生成,形成类级别特征,并在训练中动态调整,增强几何理解能力。同时,设计基于柯尔莫哥洛夫-阿诺德网络(KAN)的混合解码器,克服传统线性解码器对复杂多模态输入建模不足的问题。在Pix3D数据集上的大量实验表明,MGP-KAD达到当前最优(SOTA)表现,显著提升几何完整性、表面光滑度与细节保真度。本工作为复杂场景下的单视角3D重建提供了鲁棒有效的解决方案。

原文摘要 · Abstract (English)

Single-view 3D reconstruction in complex real-world scenes is challenging due to noise, object diversity, and limited dataset availability. To address these challenges, we propose MGP-KAD, a novel multimodal feature fusion framework that integrates RGB and geometric prior to enhance reconstruction accuracy. The geometric prior is generated by sampling and clustering ground-truth object data, producing class-level features that dynamically adjust during training to improve geometric understanding. Additionally, we introduce a hybrid decoder based on Kolmogorov-Arnold Networks (KAN) to overcome the limitations of traditional linear decoders in processing complex multimodal inputs. Extensive experiments on the Pix3D dataset demonstrate that MGP-KAD achieves state-of-the-art (SOTA) performance, significantly improving geometric integrity, smoothness, and detail preservation. Our work provides a robust and effective solution for advancing single-view 3D reconstruction in complex scenes.

3D重建多模态融合几何先验KAN

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。