arXiv:2510.22682eess.AS2025-10被引 2

让语音定位模型自己判断结果靠不靠谱,提升复杂环境下的定位精度。

SRP-PHAT-NET: A Reliability-Driven DNN for Reverberant Speaker Localization

  • 用SRP-PHAT方向图作输入,训练时用高斯加权标签增强可靠性评估
  • 通过调节高斯核宽度,可按需平衡定位准确率与置信度得分
  • 只采纳高置信度预测,定位误差显著降低,适合实际部署场景

在混响环境中精确估计声源方向(DOA)仍是空间音频应用的关键挑战。尽管深度学习方法在该条件下表现优异,但通常缺乏对预测结果可靠性的评估机制——这对真实应用至关重要。本文提出SRP-PHAT-NET,一种基于SRP-PHAT方向图作为空间特征的深度神经网络框架,并引入内置的可靠性估计。为实现有意义的可靠性评分,模型采用围绕真实方向的高斯加权标签进行训练。我们系统分析了标签平滑对准确率与可靠性的影响,表明高斯核宽度的选择可针对具体应用场景进行调优。实验结果表明,仅使用高置信度预测能显著提升定位精度,凸显将可靠性集成到深度学习型DOA估计中的实用价值。

原文摘要 · Abstract (English)

Accurate Direction-of-Arrival (DOA) estimation in reverberant environments remains a fundamental challenge for spatial audio applications. While deep learning methods have shown strong performance in such conditions, they typically lack a mechanism to assess the reliability of their predictions - an essential feature for real-world deployment. In this work, we present the SRP-PHAT-NET, a deep neural network framework that leverages SRP-PHAT directional maps as spatial features and introduces a built-in reliability estimation. To enable meaningful reliability scoring, the model is trained using Gaussian-weighted labels centered around the true direction. We systematically analyze the influence of label smoothing on accuracy and reliability, demonstrating that the choice of Gaussian kernel width can be tuned to application-specific requirements. Experimental results show that selectively using high-confidence predictions yields significantly improved localization accuracy, highlighting the practical benefits of integrating reliability into deep learning-based DOA estimation.

语音定位深度学习可靠性评估混响环境

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。