首个针对非语言发声的标准化评估基准,提升语音合成表达力评测可信度。
NV-Bench: Benchmark of Nonverbal Vocalization Synthesis for Expressive Text-to-Speech Generation
- 按交际功能分类14类非语言发声,构建多语言真实语料库
- 提出双维度评估:可控性(PCER)与音质真实性(分布差距)
- 验证客观指标与人类感知高度相关,适合研究表达式语音合成者
尽管近期文本到语音(TTS)系统越来越多地整合非语言发声(NVs),但其评估仍缺乏标准指标和可靠参照。为此,我们提出NV-Bench,首个基于功能分类体系的基准,将非语言发声视为交际行为而非声学现象。NV-Bench包含1,651条多语言、真实场景的语句,每条配有真人参考音频,覆盖14类非语言发声且类别均衡。我们引入双维度评估协议:(1) 指令对齐,采用提出的副语言字符错误率(PCER)评估可控性;(2) 声学保真度,通过与真实录音的分布差距衡量音质真实感。我们评估了多种TTS模型并建立了两个基线。实验结果表明,我们的客观指标与人类感知高度相关,确立了NV-Bench作为标准化评估框架的地位。
原文摘要 · Abstract (English)
While recent text-to-speech (TTS) systems increasingly integrate nonverbal vocalizations (NVs), their evaluations lack standardized metrics and reliable ground-truth references. To bridge this gap, we propose NV-Bench, the first benchmark grounded in a functional taxonomy that treats NVs as communicative acts rather than acoustic artifacts. NV-Bench comprises 1,651 multi-lingual, in-the-wild utterances with paired human reference audio, balanced across 14 NV categories. We introduce a dual-dimensional evaluation protocol: (1) Instruction Alignment, utilizing the proposed paralinguistic character error rate (PCER) to assess controllability, (2) Acoustic Fidelity, measuring the distributional gap to real recordings to assess acoustic realism. We evaluate diverse TTS models and develop two baselines. Experimental results demonstrate a strong correlation between our objective metrics and human perception, establishing NV-Bench as a standardized evaluation framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。