arXiv:2605.03934cs.SDcs.AI2026-05中稿 · Signal Processing被引 1

让声音检测模型能识别新声音并持续学习,更贴近真实环境。

Towards Open World Sound Event Detection

论文配图:Towards Open World Sound Event Detection
图 1 · 摘自论文原文
  • 用可变形注意力聚焦关键时间片段,自适应捕捉声音特征。
  • 在开放世界场景下,检测新声音的能力显著优于已有方法。
  • 适合需要持续学习新声音的应用,如智能城市与医疗监测。

声音事件检测(SED)在监控、智慧城市、医疗和多媒体索引中至关重要。但传统SED系统基于封闭世界假设,在真实环境中面对新声音时表现受限。受计算机视觉中开放世界学习的启发,本文提出开放世界声音事件检测(OW-SED)范式:模型需识别已知事件、发现未知事件,并增量学习。针对重叠与模糊事件等挑战,提出1D可变形架构,利用可变形注意力自适应聚焦显著时间区域。进一步设计开放世界可变形声音事件检测变换器(WOOT),包含特征解耦机制分离类别相关与无关表示,结合一对多匹配策略和多样性损失以增强表征多样性。实验表明,该方法在封闭世界设置下性能略优,而在开放世界场景下显著超越现有基线。

原文摘要 · Abstract (English)

Sound Event Detection (SED) plays a vital role in audio understanding, with applications in surveillance, smart cities, healthcare, and multimedia indexing. However, conventional SED systems operate under a closed-world assumption, limiting their effectiveness in real-world environments where novel acoustic events frequently emerge. Inspired by the success of open-world learning in computer vision, we introduce the Open-World Sound Event Detection (OW-SED) paradigm, where models must detect known events, identify unseen ones, and incrementally learn from them. To tackle the unique challenges of OW-SED, such as overlapping and ambiguous events, we propose a 1D Deformable architecture that leverages deformable attention to adaptively focus on salient temporal regions. Furthermore, we design a novel Open-World Deformable Sound Event Detection Transformer (WOOT) framework incorporating feature disentanglement to separate class-specific and class-agnostic representations, together with a one-to-many matching strategy and a diversity loss to enhance representation diversity. Experimental results demonstrate that our method achieves marginally superior performance compared to existing leading techniques in closed-world settings and significantly improves over existing baselines in open-world scenarios.

声音检测开放世界可变形注意力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。