构建首个细粒度语音事件识别基准,精准定位笑声哭声等非语言声音。
WESR: Scaling and Evaluating Word-level Event-Speech Recognition
- 定义21类语音事件,区分独立与混合于语音的类型。
- 建立900+条专家标注数据集,支持精确定位评估。
- 训练超1700小时专用模型,性能超越开源与商用方案。
语音不仅传递语言信息,还包含丰富的非语言声音事件,如笑、哭等。尽管语义转录研究成熟,但非语言事件的精确定位仍是关键且未充分探索的挑战。现有方法存在任务定义不明确、类别覆盖有限、时间粒度模糊等问题,且缺乏标准化评估框架,制约下游应用发展。为此,我们首先构建了包含21类语音事件的精细分类体系,将事件分为离散(独立)与连续(与语音混合)两类。基于该分类体系,提出WESR-Bench——一个专家标注的评估数据集(900+个语音片段),采用新型位置感知协议,分离语音识别误差与事件检测误差,实现对离散与连续事件的精确时空定位测量。同时,构建超过1700小时的专用语料库,训练专用模型,在保持高质量语音识别的同时,显著优于开源音频-语言模型及商用API。我们期待WESR成为未来复杂真实听觉场景建模的重要基础资源。
原文摘要 · Abstract (English)
Speech conveys not only linguistic information but also rich non-verbal vocal events such as laughing and crying. While semantic transcription is well-studied, the precise localization of non-verbal events remains a critical yet under-explored challenge. Current methods suffer from insufficient task definitions with limited category coverage and ambiguous temporal granularity. They also lack standardized evaluation frameworks, hindering the development of downstream applications. To bridge this gap, we first develop a refined taxonomy of 21 vocal events, with a new categorization into discrete (standalone) versus continuous (mixed with speech) types. Based on the refined taxonomy, we introduce WESR-Bench, an expert-annotated evaluation set (900+ utterances) with a novel position-aware protocol that disentangles ASR errors from event detection, enabling precise localization measurement for both discrete and continuous events. We also build a strong baseline by constructing a 1,700+ hour corpus, and train specialized models, surpassing both open-source audio-language models and commercial APIs while preserving ASR quality. We anticipate that WESR will serve as a foundational resource for future research in modeling rich, real-world auditory scenes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。