arXiv:2602.06271cs.SDeess.AS2026-02

用合成音景训练模型,检测引发烦躁的特定声音。

Misophonia Trigger Sound Detection on Synthetic Soundscapes Using a Hybrid Model with a Frozen Pre-Trained CNN and a Time-Series Module

  • 用音频合成生成模拟患者触发声的音景数据。
  • 双向门控循环单元模型在多类检测中表现最佳。
  • 轻量级网络适合少量样本个性化适配,实用性强。

Misophonia是一种对特定日常声音(触发声)耐受性降低的障碍,可能引发愤怒、恐慌或焦虑等强烈负面情绪,严重影响日常生活与生活质量。可协助识别触发声的技术有望减轻痛苦、提升福祉。本研究探索声音事件检测(SED),以定位连续环境音频中的触发声时间段,作为辅助支持的基础步骤。由于真实世界中的misophonia数据稀缺,我们采用音频合成技术生成专为触发声检测定制的合成音景。随后,使用融合预训练冻结CNN与可训练时序模块(如GRU、LSTM、ESN及其双向变体)的混合模型进行检测任务。性能通过常见SED指标评估,包括多音检测得分1(PSDS1)。在多类触发声检测任务中,双向时序建模持续提升性能,其中双向门控循环单元(BiGRU)达到最佳准确率。值得注意的是,双向回声状态网络(BiESN)仅优化读出层,参数量少几个数量级,仍保持竞争力。进一步通过最多5个支持样本的“进食声”小样本检测任务模拟用户个性化,结果显示BiGRU与BiESN均表现稳健,表明轻量化时序模块在个性化触发声检测中具有前景。

原文摘要 · Abstract (English)

Misophonia is a disorder characterized by a decreased tolerance to specific everyday sounds (trigger sounds) that can evoke intense negative emotional responses such as anger, panic, or anxiety. These reactions can substantially impair daily functioning and quality of life. Assistive technologies that selectively detect trigger sounds could help reduce distress and improve well-being. In this study, we investigate sound event detection (SED) to localize intervals of trigger sounds in continuous environmental audio as a foundational step toward such assistive support. Motivated by the scarcity of real-world misophonia data, we generate synthetic soundscapes tailored to misophonia trigger sound detection using audio synthesis techniques. Then, we perform trigger sound detection tasks using hybrid CNN-based models. The models combine feature extraction using a frozen pre-trained CNN backbone with a trainable time-series module such as gated recurrent units (GRUs), long short-term memories (LSTMs), echo state networks (ESNs), and their bidirectional variants. The detection performance is evaluated using common SED metrics, including Polyphonic Sound Detection Score 1 (PSDS1). On the multi-class trigger SED task, bidirectional temporal modeling consistently improves detection performance, with Bidirectional GRU (BiGRU) achieving the best overall accuracy. Notably, the Bidirectional ESN (BiESN) attains competitive performance while requiring orders of magnitude fewer trainable parameters by optimizing only the readout. We further simulate user personalization via a few-shot "eating sound" detection task with at most five support clips, in which BiGRU and BiESN are compared. In this strict adaptation setting, BiESN shows robust and stable performance, suggesting that lightweight temporal modules are promising for personalized misophonia trigger SED.

声音检测轻量化模型个性化合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。