arXiv:2604.16211cs.SD2026-04中稿 · as a long paper at…被引 1

评测语音生成中非语言发声的准确性与自然度

NVV-SuperBench: Beyond Words, Beyond Quality-Benchmarking Nonverbal Vocalizations in Speech Generation

论文配图:NVV-SuperBench: Beyond Words, Beyond Quality-Benchmarking Nonverbal Vocalizations in Speech Generation
图 1 · 摘自论文原文
  • 构建中英双语非语言发声评测基准,覆盖45类发声类型
  • 发现语音质量与发声控制能力常不匹配,长时情感发声难处理
  • 适合语音合成、人机交互研究者参考

非语言发声(NVVs),如笑声、叹息和啜泣,是实现自然语音的关键,但现有评估标准很少同时衡量系统是否生成了预期的发声、位置是否准确以及发声是否清晰。本文提出NVV-SuperBench,一个涵盖中英文的语音生成非语言发声评测基准。该基准提供统一的45类发声分类体系和多维度评估协议,超越传统语音质量评估,重点考察发声的可控性、位置准确性和感知显著性。我们对15种语音生成系统进行了评测,涵盖基于提示和基于标签的控制范式,采用客观指标、人工听感测试及大模型多评分器评估。结果表明,发声可控性常与语音质量脱钩,低信噪比口腔线索和长时间情感发声仍是主要瓶颈。NVV-SuperBench揭示了当前技术差距,助力更拟人化语音生成的发展。

原文摘要 · Abstract (English)

Nonverbal vocalizations (NVVs), such as laughing, sighing, and sobbing, are essential for human-like speech, yet standardized evaluation rarely jointly assesses whether systems generate the intended NVVs, place them correctly, and keep them salient without harming speech. We present NVV-SuperBench, a bilingual English/Chinese benchmark for speech generation with NVVs. It provides a unified 45-type taxonomy and a multi-axis protocol beyond conventional speech quality assessment, evaluating NVV-specific controllability, placement, and perceptual salience. We benchmark 15 speech generation systems spanning prompt-based and tag-based control paradigms, using objective metrics, human listening tests, and LLM-based multi-rater evaluation. Results show that NVV controllability often decouples from speech quality, while low-SNR oral cues and long-duration affective NVVs remain bottlenecks. NVV-SuperBench highlights current gaps and supports progress toward more human-like speech generation.

语音生成非语言发声评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。