arXiv:2507.09606cs.SDeess.AS2025-07

用集成方法降低声音检测模型在开放环境中的过度自信

Ensemble Confidence Calibration for Sound Event Detection in Open-environment

  • 采用集成方法结合能量感知的软最大值校准
  • 显著降低对未知场景的过度置信,提升泛化能力
  • 适合需要可靠不确定度估计的现实音频应用

声音事件检测(SED)在受控环境中已取得显著进展,但真实应用场景多处于开放环境。此时,现有方法常产生过高的预测置信度,且缺乏有效的不确定性度量方式,限制其在新场景下的适应能力。为此,本文首次在SED中引入集成方法以增强对域外输入(OOD)的鲁棒性。提出一种基于能量的开放世界软最大值校准方法(EOW-Softmax),使系统能更好处理未知场景中的不确定性。进一步将该方法应用于声音发生与重叠检测(SOD),通过调整预测结果,在保持重叠事件检测能力的同时提升模型适应性。实验表明,该方法有效降低了过度自信,增强了对域外情况的应对能力。

原文摘要 · Abstract (English)

Sound event detection (SED) has made strong progress in controlled environments with clear event categories. However, real-world applications often take place in open environments. In such cases, current methods often produce predictions with too much confidence and lack proper ways to measure uncertainty. This limits their ability to adapt and perform well in new situations. To solve this problem, we are the first to use ensemble methods in SED to improve robustness against out-of-domain (OOD) inputs. We propose a confidence calibration method called Energy-based Open-World Softmax (EOW-Softmax), which helps the system better handle uncertainty in unknown scenes. We further apply EOW-Softmax to sound occurrence and overlap detection (SOD) by adjusting the prediction. In this way, the model becomes more adaptable while keeping its ability to detect overlapping events. Experiments show that our method improves performance in open environments. It reduces overconfidence and increases the ability to handle OOD situations.

声音检测不确定性开放世界集成学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。