仅用声音就能准确估算多人3D姿态,突破声学信号重叠难题。
Sound-based Multi-Person 3D Pose Estimation

- 设计多尺度声学编码器分离重叠声纹,捕捉细微运动特征。
- 引入时序姿态解码器,通过注意力机制解析多人动态关系。
- 构建6小时43.2万帧的AMP数据集,验证方法有效性。
能否仅通过声音恢复多人的3D姿态?本文首次尝试仅基于声学信号进行多人群体3D姿态估计。由于运动相关的信号变化相互叠加,该任务极具挑战性:多人存在导致声学特征重叠,难以将特定信号变化归因于个体姿态;此外,人与人之间的反射造成复杂传播延迟,模糊了时间上运动与声学间的关联。为此,我们提出SoundMHPE(基于声音的多人群体姿态估计器),一个包含两个关键组件的编码器-解码器框架。首先,声学多尺度编码器捕捉多样化的时序与精细频率特征,从复杂重叠信号中分离出微弱声学线索;其次,时序姿态解码器采用注意力机制,在连续帧间解耦多人信息,联合考虑时序动态与人际依赖,实现逐帧个体姿态的精准重建。为验证方法,我们构建了6小时、包含43.2万帧同步多人群体姿态与声学数据的AMP数据集,并证明SoundMHPE优于基线模型。
原文摘要 · Abstract (English)
Can we recover the 3D poses of multiple people using only sound? This paper presents the first attempt to estimate multi-person 3D poses solely from acoustic signals. Estimating the poses of multiple individuals using acoustic signals is inherently challenging due to the superposition of motion-dependent signal variations. Unlike single-person scenarios, the presence of multiple subjects leads to overlapping acoustic signatures, making it difficult to attribute specific signal changes to an individual's pose. Furthermore, the complexity is compounded by inter-person reflections, which introduce intricate propagation delays that obscure the temporal motion-acoustic relationship. To address these issues, we propose SoundMHPE (Sound-based Multi-person Human Pose Estimator), a novel encoder-decoder framework consisting of two key components. First, the Acoustic Multi-scale Encoder captures diverse temporal and fine-grained frequency features to isolate subtle acoustic signatures from complex, overlapping signals. Second, the Temporal Pose Decoder employs an attention mechanism to disentangle multi-person information across successive frames. By jointly accounting for temporal dynamics and inter-person dependencies, this component precisely reconstructs frame-wise individual poses. To validate our approach, we constructed the 6-hour Acoustic Multi-person Pose (AMP) dataset consisting of 432K synchronized frames of multi-person pose and acoustic data, and demonstrated that our SoundMHPE outperforms baseline models. Project page: https://oumi03.github.io/sound-mhpe/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。