为多说话人声源定位提供置信度量化,提升实际应用中的可靠性。
Uncertainty Quantification and Risk Control for Multi-Speaker Sound Source Localization
- 基于分位数预测框架,构建覆盖真实声源位置的置信区域。
- 在已知与未知声源数量下均实现可靠有限样本保证。
- 适合对定位可靠性要求高的智能音箱、语音会议系统等场景。
可靠的声源定位(SSL)在诸多下游任务中至关重要,其决策不仅依赖于定位精度,还需估计结果的置信度。在混响环境和多源场景等挑战条件下,这一需求尤为突出。然而,现有SSL方法通常仅提供点估计,缺乏不确定性量化(UQ)。本文利用分位数预测(Conformal Prediction, CP)框架及其扩展,针对一般风险函数控制,提出两种互补的UQ方法:第一种假设活跃声源数量已知,构建覆盖真实声源位置的预测区域;第二种处理更复杂的声源数量未知情形,先可靠估计声源数量,再形成对应预测区域。我们在不同混响水平和源配置下的大量仿真与真实录音上进行了评估,结果表明,两种方法在已知与未知源数场景下均具备可靠的有限样本保证,且性能稳定,凸显了所提框架在不确定性感知的声源定位中的实用价值。
原文摘要 · Abstract (English)
Reliable Sound Source Localization (SSL) plays an essential role in many downstream tasks, where informed decision making depends not only on accurate localization but also on the confidence in each estimate. This need for reliability becomes even more pronounced in challenging conditions, such as reverberant environments and multi-source scenarios. However, existing SSL methods typically provide only point estimates, offering limited or no Uncertainty Quantification (UQ). We leverage the Conformal Prediction (CP) framework and its extensions for controlling general risk functions to develop two complementary UQ approaches for SSL. The first assumes that the number of active sources is known and constructs prediction regions that cover the true source locations. The second addresses the more challenging setting where the source count is unknown, first reliably estimating the number of active sources and then forming corresponding prediction regions. We evaluate the proposed methods on extensive simulations and real-world recordings across varying reverberation levels and source configurations. Results demonstrate reliable finite-sample guarantees and consistent performance for both known and unknown source-count scenarios, highlighting the practical utility of the proposed frameworks for uncertainty-aware SSL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。