让语音模型学会自然表达笑声、嗯啊等非语言声音,提升合成语音的真实感。
NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations
- 构建包含48,430条带标注的非语言语音数据集,覆盖18类声音
- 开发可识别并合成非语言音的语音模型,支持在任意位置插入语气词
- 首个大规模中文非语言语音数据集,适合做语音交互与情感合成研究
非语言语音(如笑声、呼吸声、'嗯'、'哦'等)在自然口语交流中至关重要,但传统自动语音识别(ASR)和文本转语音(TTS)系统常忽略此类线索。本文提出NVSpeech,一个集成且可扩展的语音建模流程,涵盖数据构建、ASR建模与可控TTS。首先,构建了48,430条人工标注的人类语句数据集,包含18类词级非语言语音类别。其次,设计非语言感知的ASR模型,将非语言信号作为可解码的内嵌标记(如“你真有趣[笑声]”),实现词汇与非语言信息的联合转录,并据此自动标注出首个大规模中文数据集——174,179条语句(573小时),含词级对齐与非语言标记。最后,在人工与自标注数据上微调零样本TTS模型,实现对非语言语音的显式控制,可在任意词元位置进行上下文感知插入,生成更自然的语音。NVSpeech首次提供开放、大规模、词级标注的中文情感语音建模管道,实现了识别与生成的统一与可控。数据集与音频演示见https://nvspeech170k.github.io/。
原文摘要 · Abstract (English)
Paralinguistic vocalizations-including non-verbal sounds like laughter and breathing, as well as lexicalized interjections such as "uhm" and "oh"-are integral to natural spoken communication. Despite their importance in conveying affect, intent, and interactional cues, such cues remain largely overlooked in conventional automatic speech recognition (ASR) and text-to-speech (TTS) systems. We present NVSpeech, an integrated and scalable pipeline that bridges the recognition and synthesis of paralinguistic vocalizations, encompassing dataset construction, ASR modeling, and controllable TTS. (1) We introduce a manually annotated dataset of 48,430 human-spoken utterances with 18 word-level paralinguistic categories. (2) We develop the paralinguistic-aware ASR model, which treats paralinguistic cues as inline decodable tokens (e.g., "You're so funny [Laughter]"), enabling joint lexical and non-verbal transcription. This model is then used to automatically annotate a large corpus, the first large-scale Chinese dataset of 174,179 utterances (573 hours) with word-level alignment and paralingustic cues. (3) We finetune zero-shot TTS models on both human- and auto-labeled data to enable explicit control over paralinguistic vocalizations, allowing context-aware insertion at arbitrary token positions for human-like speech synthesis. By unifying the recognition and generation of paralinguistic vocalizations, NVSpeech offers the first open, large-scale, word-level annotated pipeline for expressive speech modeling in Mandarin, integrating recognition and synthesis in a scalable and controllable manner. Dataset and audio demos are available at https://nvspeech170k.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。