arXiv:2606.06921cs.SD2026-06中稿 · Interspeech 2026

构建新数据集测试声音场景分类模型抗干扰能力

Towards Event-Robust Acoustic Scene Classification

论文配图:Towards Event-Robust Acoustic Scene Classification
图 1 · 摘自论文原文
  • 用大模型在背景音中插入突发声事件模拟真实环境
  • 现有模型在新数据集上性能大幅下降,最高降37%
  • 适合研究鲁棒性、实际应用中的音频系统优化

本文提出事件偏移声音场景(ESAS)数据集,用于评估声音场景分类(ASC)系统对未知声事件的鲁棒性。现有ASC数据集多为干净一致的音频,而真实环境常出现多样且意外的声音事件。为弥合这一差距,ESAS通过大语言模型辅助,在背景场景中注入前景声事件,模拟真实声学变化。本文介绍数据集构建方法、统计信息及评测协议,并对当前主流ASC系统进行评估。实验表明,现有模型在事件偏移挑战下性能显著下降,最高降幅达37%。该数据集旨在推动未来研究向事件鲁棒的ASC方向发展。

原文摘要 · Abstract (English)

This paper introduces the Event-Shifted Acoustic Scene (ESAS) dataset, a novel benchmark for evaluating the robustness of Acoustic Scene Classification (ASC) systems against unknown sound events. Existing ASC datasets typically contain recordings of clean and consistent audio, while real-world environments often include diverse and unexpected sound events. To bridge this gap, ESAS simulates real-world acoustic variability by injecting foreground sound events into background scenes with the assistance of large language models. In this work, we present the construction methodology, dataset statistics, and evaluation protocols. Furthermore, a comprehensive evaluation of state-of-the-art ASC systems is conducted using the ESAS benchmark. Experimental results reveal that existing ASC models suffer significant performance degradation when facing the event-shift challenge. The introduction of the ESAS dataset aims to drive future research toward event-robust ASC.

声音分类鲁棒性数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。