arXiv:2606.15888cs.SDcs.AI2026-06

首个能准确评估语音中非语言发声质量的模型。

NVMOS: Non-Verbal Vocalization Quality Assessment in Speech

论文配图:NVMOS: Non-Verbal Vocalization Quality Assessment in Speech
图 1 · 摘自论文原文
  • 设计局部发声事件聚焦模块,精准捕捉非语言发声特征。
  • 在专家评分上达到或超过人类水平的一致性。
  • 揭示通用多模态大模型在该任务上的不可靠性,适合语音合成研究者。

非语言发声(NVs),如笑声、叹气和咳嗽,是情感与意图的重要声学线索。现有语音质量评估方法多关注整体自然度,而针对非语言语音合成(NV-TTS)的评价主要关注目标发声类型和位置是否正确,但对发声本身感知质量的研究仍不足。为此,我们构建了包含多个NV-TTS系统输出及真实发声样本的NV-MOS数据集,并由三位声学专家在感知质量量表上进行评分。进一步分析发现,像Gemini这样的多模态大语言模型在评分上与专家意见存在明显不一致。结果表明,通用多模态模型无法可靠替代人工判断。因此,我们提出NVMOS,据知是首个能可靠预测语音中非语言发声感知质量的模型。实验显示,通过引入局部发声事件聚焦模块,NVMOS在与人类专家评分的一致性上达到或超越专家水平。

原文摘要 · Abstract (English)

Non-verbal vocalizations (NVs), such as laughter, sighs, and coughs, are important acoustic cues for emotion and intent. Existing speech quality assessment methods typically focus on overall naturalness, while non-verbal TTS evaluations mainly examine whether a target NV appears with the correct type and position. However, the perceptual quality of NV events themselves remains underexplored. To address this gap, we construct an NV-MOS dataset containing outputs from multiple NV-TTS systems and naturally occurring NV samples, with ratings collected from three acoustic experts on a perceptual quality scale. We further analyze audio-capable multimodal large language models such as Gemini and find clear inconsistencies between their scores and expert ratings. These results suggest that general-purpose multimodal models cannot reliably replace human judgments for NV quality assessment. We then propose NVMOS, to our knowledge the first model that can reliably predict the perceptual quality of NV events in speech. Experimental results show that, with a local NV-event focusing module, NVMOS reaches expert-level or stronger agreement with human MOS.

语音合成质量评估非语言发声多模态模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。