arXiv:2607.01563eess.AS2026-07

让语音识别系统学会识别笑声、咳嗽等非语言声音,提升对话情感理解能力。

Beyond Words: Towards Effective Modeling of Non-Verbal Vocalizations in ASR

论文配图:Beyond Words: Towards Effective Modeling of Non-Verbal Vocalizations in ASR
图 1 · 摘自论文原文
  • 分两阶段训练:先统一标记所有非语言声,再针对性微调
  • 利用高频声音(如笑)的知识帮助识别罕见声音(如哭)
  • 通过声音转换增强数据平衡,显著提升稀有类别识别率

现代自动语音识别(ASR)系统在转录词汇内容方面表现优异,但常忽略笑声、呼吸、咳嗽、哭泣等非语言发声(NVs),这些发声承载着对话与情感信息。由于NV标注稀疏且分布极不均衡(如呼吸、笑声频繁,而哭泣、咳嗽罕见),建模困难。本文研究三种数据驱动策略:(1)两阶段课程学习——先将所有NV事件映射为通用标记,再在目标类别上微调;(2)高资源事件(如笑声、呼吸)向低资源事件(如哭泣)进行跨标记迁移;(3)基于语音转换的增强方法结合类别平衡。实验表明,可利用各类发声间的共享声学结构,有效提升稀有类别的检测性能,同时保持原有词汇识别质量。

原文摘要 · Abstract (English)

Modern automatic speech recognition (ASR) systems excel at transcribing lexical content but often omit nonverbal vocalizations (NVs), such as laughter, breaths, coughs, and cries, that carry conversational and affective information. Modeling NVs in ASR is challenging because NV annotations are sparse and highly long-tailed, with frequent categories such as breaths and laughter dominating rarer events such as cries and coughs. We study three data-centric strategies for improving low-resource NV recognition: (1) a two-stage curriculum that first maps all NV events to a generic token and then fine-tunes on target categories; (2) inter-token transfer from high-resource events, such as laughter and breath, to rare events, such as crying; and (3) voice-conversion augmentation with class balancing. Experiments show that shared acoustic structure across vocal events can be exploited to improve rare-category detection while preserving lexical ASR quality.

语音识别非语言声数据增强少样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。