用声音估计3D人体姿态,不受人站位影响。
Acoustic-based 3D Human Pose Estimation Robust to Human Position
- 用对抗学习提取与位置无关的声学特征
- 利用前置信号作参考,抗反射和衍射干扰
- 新数据集验证,在多位置下性能更优
本文研究仅从低级声学信号进行3D人体姿态估计的问题。现有基于主动声学感知的方法隐含假设目标用户位于扬声器与麦克风之间的直线上。由于人体对声音的反射和衍射会导致声信号细微变化(相比遮挡),当受试者偏离该直线时,现有模型精度显著下降,限制了其在真实场景中的实用性。为此,我们提出一种新方法,包含位置判别器和抗混响模型:前者通过对抗学习提取位置不变特征;后者利用估计目标时刻之前的声信号作为参考,增强对因衍射和反射导致的声音到达时间变化的鲁棒性。我们构建了一个覆盖多种人体位置的声学姿态估计数据集,并通过实验表明,所提方法优于现有方法。
原文摘要 · Abstract (English)
This paper explores the problem of 3D human pose estimation from only low-level acoustic signals. The existing active acoustic sensing-based approach for 3D human pose estimation implicitly assumes that the target user is positioned along a line between loudspeakers and a microphone. Because reflection and diffraction of sound by the human body cause subtle acoustic signal changes compared to sound obstruction, the existing model degrades its accuracy significantly when subjects deviate from this line, limiting its practicality in real-world scenarios. To overcome this limitation, we propose a novel method composed of a position discriminator and reverberation-resistant model. The former predicts the standing positions of subjects and applies adversarial learning to extract subject position-invariant features. The latter utilizes acoustic signals before the estimation target time as references to enhance robustness against the variations in sound arrival times due to diffraction and reflection. We construct an acoustic pose estimation dataset that covers diverse human locations and demonstrate through experiments that our proposed method outperforms existing approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。