arXiv:2608.24163eess.AScs.AI2026-08

提出新评估指标,让语音合成更自然地表达笑声等非语言声音

Preference Optimization for Non-Verbal Vocalization Synthesis

  • 将非语言音效视为独立符号,设计新型评估指标NV-CER
  • 在18种非语言声音上验证,标准DPO方法效果最佳
  • 适合想提升语音合成情感表达力的研究者和开发者

非语言音效(如笑声、咳嗽、叹气)对富有表现力的语音合成至关重要,但其偏好优化的有效性仍不明确。本文系统研究了具备非语言音效生成能力的语音合成中的偏好优化,聚焦于偏好信号、偏好对构建及基于直接偏好优化(DPO)的目标函数。通过将非语言音效标签视为独立输出符号,提出一种非语言音效感知字符错误率(NV-CER),并在语音与非语言内容上计算加权拼音基错误率,实现无需修改底层优化算法即可控制非语言音效生成。在Emilia-NV和扩展的NV-Bench数据集(涵盖18种非语言音效类型)上的实验揭示了不同设计选择对非语言音效生成与词汇保真度的影响,并确立了使用标准DPO的有效配置。客观评估、基于大模型的评估及人工评估结果一致支持上述发现,为富有表现力语音合成的非语言音效感知后训练提供了实用指导。

原文摘要 · Abstract (English)

Non-verbal vocalizations (NVs), such as laughter, coughs, and sighs, are essential for expressive TTS, but the effectiveness of preference optimization for NV generation remains poorly understood. We systematically study preference optimization for NV-capable TTS, focusing on preference signals, preference-pair construction, and DPO-based optimization objectives. We formulate an NV-aware character error rate (NV-CER) by treating NV tags as distinct output symbols and computing a weighted pinyin-based CER over both verbal and non-verbal content, enabling controllable optimization of NV realization without modifying the underlying optimization algorithm. Experiments on Emilia-NV and the augmented NV-Bench covering 18 NV types reveal how different design choices affect NV realization and lexical fidelity, and establish an effective setup using standard DPO. Objective, LLM-based, and human evaluations provide converging evidence for our findings, offering practical insights into NV-aware post-training for expressive TTS.

语音合成非语言音效偏好优化评估指标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。