arXiv:2607.12703eess.AS2026-07中稿 · IWAENC 2026

提出音频分段新任务,自动识别未知声音的开始与结束时间。

Audio Diarization: A New Paradigm for Exploring Audio Recordings with Unknown Event Classes

  • 基于说话人分割思路,实现无提示的开放集声音事件定位。
  • 在未知声音场景下表现接近封闭集检测系统。
  • 适合智能监控、环境音频分析等需发现新声音的应用。

我们提出一项新任务——音频分段(audio diarization),旨在解决未知环境中声音事件类别未知的场景。传统方法需预先定义事件类别,而本工作聚焦于第一步:在不依赖用户提示的前提下,定位开放集中任意声音事件的起止时间,允许多个事件重叠。该任务类比于多说话人语音处理中的说话人分段阶段,为后续声音分类提供简化输入。我们展示了如何将现有说话人分段系统调整用于音频分段,并提出了相应的评估框架。实验表明,该系统在性能上与封闭集声音事件检测相当,同时具备发现新声音的能力。

原文摘要 · Abstract (English)

We propose a new task, audio diarization. The motivation is that there are applications, such as audio monitoring in an unknown environment, where initially the sound event classes to be recognized are unknown. For such a scenario, we propose to first localize in time relevant sound events and to classify them, e.g., by comparing with known event classes, in a second step. This contribution is dedicated to the first step, which we call audio diarization, as it is reminiscent of the speaker diarization stage that precedes and simplifies the second stage, speech recognition, in multi-talker conversational speech processing. In this contribution, we define audio diarization as detecting onset and offset times of sound events with overlap for an open set of classes and without user prompts. We show how a speaker diarization system can be adjusted for audio diarization and propose an evaluation setup. Compared to a closed-set sound event detection system, the proposed system achieves similar performance with the additional ability to detect novel sounds.

音频分段开放集声音事件

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。