用声音信号辅助视觉重建3D人体网格,解决暗光、遮挡等难题。
Sonicmesh: Enhancing 3D Human Mesh Reconstruction in Vision-Impaired Environments With Acoustic Signals
- 融合声学信号与RGB图像,用改进HRNet提取声纹特征。
- 在遮挡、非视线、弱光环境下重建精度显著提升。
- 适合智能监控、无障碍交互等隐私敏感场景使用。
从2D RGB图像进行3D人体网格重建(HMR)在光照不足、隐私顾虑或遮挡环境下面临挑战。声学信号具有广泛可用性、易部署且可穿透障碍物,能弥补RGB成像的不足。然而,现有方法未能有效结合声学信号与RGB数据实现鲁棒的3D HMR。主要困难在于声学生成图像分辨率低,且缺乏专用处理骨干网络。本文提出SonicMesh,通过改进现有方法HRNet以有效提取声学信号特征,并引入通用特征嵌入技术增强跨模态特征对齐精度,从而实现高精度重建。实验表明,SonicMesh在遮挡、非视线及弱光等复杂环境下均能准确重建3D人体网格。
原文摘要 · Abstract (English)
3D Human Mesh Reconstruction (HMR) from 2D RGB images faces challenges in environments with poor lighting, privacy concerns, or occlusions. These weaknesses of RGB imaging can be complemented by acoustic signals, which are widely available, easy to deploy, and capable of penetrating obstacles. However, no existing methods effectively combine acoustic signals with RGB data for robust 3D HMR. The primary challenges include the low-resolution images generated by acoustic signals and the lack of dedicated processing backbones. We introduce SonicMesh, a novel approach combining acoustic signals with RGB images to reconstruct 3D human mesh. To address the challenges of low resolution and the absence of dedicated processing backbones in images generated by acoustic signals, we modify an existing method, HRNet, for effective feature extraction. We also integrate a universal feature embedding technique to enhance the precision of cross-dimensional feature alignment, enabling SonicMesh to achieve high accuracy. Experimental results demonstrate that SonicMesh accurately reconstructs 3D human mesh in challenging environments such as occlusions, non-line-of-sight scenarios, and poor lighting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。