arXiv:2502.16298eess.AScs.SD2025-02中稿 · ICASSP 2025被引 21

首个专为非语言发声设计的通用表示模型,提升哭声等声音识别准确率。

voc2vec: A Foundation Model for Non-Verbal Vocalization

  • 基于10个开源数据集共125小时非语言音频训练,专注人声发声特征提取。
  • 在6个基准数据集上超越传统语音与音频模型,分类性能更优。
  • 适合医疗、人机交互等领域中非语言声音分析任务的研究者使用。

语音基础模型在语音任务中表现优异,但对非语言音频(如婴儿哭声)处理能力有限。现有音频基础模型也难以捕捉人类非语言发声的细微特征。本文提出 voc2vec,首个专为非语言人类发声设计的基础模型,仅使用公开开源的非语言音频数据进行训练,涵盖10个数据集,总计约125小时。实验表明,voc2vec 在非语言发声分类任务中表现优异,显著优于传统语音和音频基础模型。在六个不同基准数据集上,其性能持续超越 OpenSmile 与 emotion2vec 等强基线模型。据作者所知,voc2vec 是首个面向发声任务的通用表示模型。

原文摘要 · Abstract (English)

Speech foundation models have demonstrated exceptional capabilities in speech-related tasks. Nevertheless, these models often struggle with non-verbal audio data, such as vocalizations, baby crying, etc., which are critical for various real-world applications. Audio foundation models well handle non-speech data but also fail to capture the nuanced features of non-verbal human sounds. In this work, we aim to overcome the above shortcoming and propose a novel foundation model, termed voc2vec, specifically designed for non-verbal human data leveraging exclusively open-source non-verbal audio datasets. We employ a collection of 10 datasets covering around 125 hours of non-verbal audio. Experimental results prove that voc2vec is effective in non-verbal vocalization classification, and it outperforms conventional speech and audio foundation models. Moreover, voc2vec consistently outperforms strong baselines, namely OpenSmile and emotion2vec, on six different benchmark datasets. To the best of the authors' knowledge, voc2vec is the first universal representation model for vocalization tasks.

非语言发声基础模型音频表征多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。