arXiv:2607.02129cs.SD2026-07中稿 · EUSIPCO 2026

用单麦克风阵列通过相位谱图估计说话人头部朝向

Speaker head orientation estimation with a single microphone array using phase spectrogram features

论文配图:Speaker head orientation estimation with a single microphone array using phase spectrogram features
图 1 · 摘自论文原文
  • 利用短时傅里叶变换的相位信息作为深度网络输入
  • 在干净和嘈杂环境下均达到顶尖准确率,平均误差11.3度
  • 支持个性化适配,适合智能会议与车载监控场景

从音频中估计说话人头部朝向可为智能家居、会议系统和驾驶监控提供重要信息。本文提出一种新方法,使用单麦克风阵列的短时傅里叶变换相位成分作为深度神经网络输入,该网络融合卷积、循环和自注意力结构。与依赖物理启发手工特征或原始波形的已有方法不同,本方法能有效学习模拟数据与真实录音。模型在基于语音方向性模式生成的大规模数据集上训练,并在真实录音上微调,实现当前最优性能,在清洁与噪声条件下均优于基线。个性化实验显示显著提升,适应个体用户与环境后,平均角度误差降至11.3度。

原文摘要 · Abstract (English)

Estimating a speaker's head orientation from audio can provide valuable information in smart environments, meetings, and driver monitoring. We propose a novel approach that leverages the phase component of the short-time Fourier transform from a single microphone array as input to a deep neural network combining convolutional, recurrent, and self-attention layers. Unlike prior methods that use physics-informed handcrafted features or raw waveform inputs, our approach enables robust learning from simulated and real data. Trained on a large-scale dataset generated with voice directivity patterns and fine-tuned on real recordings, our model achieves state-of-the-art accuracy, outperforming baselines under both clean and noisy conditions. Personalization experiments further demonstrate significant gains, reaching a mean angular error of 11.3 degrees when adapting to individual users and environments.

声源定位头部朝向深度学习麦克风阵列

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。