arXiv:2501.05264cs.CVcs.AI2025-01中稿 · CVPR被引 11

解决多模态3D人体姿态估计中的信息不平衡问题

Towards Balanced Multi-Modal Learning in 3D Human Pose Estimation

  • 用谢林值评估各传感器贡献,识别模态不平衡
  • 设计早期学习减速策略,平衡多模态学习进度
  • 在MM-Fi数据集上验证,复杂环境下性能更优

3D人体姿态估计(3D HPE)已成为重要研究方向,尤其以基于RGB的方法为主。然而,RGB图像常受遮挡和隐私限制。因此,采用非侵入式传感器的多模态感知日益受到关注。但多模态3D HPE仍面临模态不平衡问题。本文提出一种新型平衡多模态学习方法,融合RGB、LiDAR、mmWave和WiFi数据。首先,设计基于谢林值的贡献评估算法,识别模态间不平衡;其次,提出模态学习调节策略,在训练初期减缓学习速度以实现平衡。在广泛使用的多模态数据集MM-Fi上进行大量实验,结果表明该方法在复杂条件下显著提升3D姿态估计性能。代码已开源:https://github.com/MICLAB-BUPT/AWC。

原文摘要 · Abstract (English)

3D human pose estimation (3D HPE) has emerged as a prominent research topic, particularly in the realm of RGB-based methods. However, the use of RGB images is often limited by issues such as occlusion and privacy constraints. Consequently, multi-modal sensing, which leverages non-intrusive sensors, is gaining increasing attention. Nevertheless, multi-modal 3D HPE still faces challenges, including modality imbalance. In this work, we introduce a novel balanced multi-modal learning method for 3D HPE, which harnesses the power of RGB, LiDAR, mmWave, and WiFi. Specifically, we propose a Shapley value-based contribution algorithm to assess the contribution of each modality and detect modality imbalance. To address this imbalance, we design a modality learning regulation strategy that decelerates the learning process during the early stages of training. We conduct extensive experiments on the widely adopted multi-modal dataset, MM-Fi, demonstrating the superiority of our approach in enhancing 3D pose estimation under complex conditions. Our source code is available at https://github.com/MICLAB-BUPT/AWC.

3D姿态估计多模态学习传感器融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。