arXiv:2409.04851cs.CV2024-09被引 6

自适应融合框架,支持任意传感器组合重建3D人体

AdaptiveFusion: Adaptive Multi-Modal Multi-View Fusion for 3D Human Body Reconstruction

  • 用Transformer统一处理多模态多视角输入,无需定制设计
  • 单个模型可处理任意数量输入,对噪声模态和未标定设备鲁棒
  • 在多个大规模数据集上优于现有方法,适合复杂环境应用

传感器技术和深度学习的进步推动了3D人体重建的发展。然而,现有方法通常依赖特定传感器,易受单一传感模态固有局限影响;且多模态融合方法需针对不同传感器组合定制设计,缺乏通用性。此外,传统基于点图投影或Transformer的融合网络易受噪声模态和传感器姿态干扰。为解决这些问题,本文提出AdaptiveFusion,一个通用的自适应多模态多视角融合框架,可有效整合任意未标定传感器输入。通过将不同视角的各模态视为等价令牌,并利用Transformer架构的灵活性设计手工模态采样模块,AdaptiveFusion仅需单一训练网络即可应对任意数量输入,具备对噪声模态的强鲁棒性。大规模人体数据集上的实验表明,该方法在多种环境下均能实现高质量3D人体重建,且精度优于当前最优融合方法。

原文摘要 · Abstract (English)

Recent advancements in sensor technology and deep learning have led to significant progress in 3D human body reconstruction. However, most existing approaches rely on data from a specific sensor, which can be unreliable due to the inherent limitations of individual sensing modalities. Additionally, existing multi-modal fusion methods generally require customized designs based on the specific sensor combinations or setups, which limits the flexibility and generality of these methods. Furthermore, conventional point-image projection-based and Transformer-based fusion networks are susceptible to the influence of noisy modalities and sensor poses. To address these limitations and achieve robust 3D human body reconstruction in various conditions, we propose AdaptiveFusion, a generic adaptive multi-modal multi-view fusion framework that can effectively incorporate arbitrary combinations of uncalibrated sensor inputs. By treating different modalities from various viewpoints as equal tokens, and our handcrafted modality sampling module by leveraging the inherent flexibility of Transformer models, AdaptiveFusion is able to cope with arbitrary numbers of inputs and accommodate noisy modalities with only a single training network. Extensive experiments on large-scale human datasets demonstrate the effectiveness of AdaptiveFusion in achieving high-quality 3D human body reconstruction in various environments. In addition, our method achieves superior accuracy compared to state-of-the-art fusion methods.

3D重建多模态融合Transformer人体建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。