arXiv:2502.02950eess.AScs.SD2025-02中稿 · IEEE TASLP被引 5

针对语音合成中的局部问题,精细优化提升零样本系统表现

Fine-grained Preference Optimization Improves Zero-shot Text-to-Speech

  • 按音频片段细粒度标注问题类型,分组设计训练损失
  • 显著降低错误率,提升语音可懂度,减少不良样本比例
  • 数据效率更高,用更少样本达到相近效果,适合资源受限场景

将人类反馈融入文本到语音(TTS)系统以对齐人类偏好,已被证明能有效提升基于语言模型的TTS系统的鲁棒性。现有方法主要依赖于在语句层面标注的偏好数据,但影响听感的问题往往仅出现在音频样本的特定片段中,其余部分生成质量良好。本文提出细粒度偏好优化方法(FPO),聚焦于解决生成样本中的局部问题,而非统一优化整个语句。我们首先分析生成样本中的问题类型,将其分为两类,并针对每类问题提出选择性训练损失策略,利用细粒度标签进行偏好优化。实验表明,FPO通过有效解决局部问题,显著提升了零样本TTS系统的鲁棒性,大幅降低错误案例比例并改善语音可懂度。此外,相较于基线系统,FPO展现出更优的数据效率,在更少训练样本下实现相似性能。

原文摘要 · Abstract (English)

Integrating human feedback to align text-to-speech (TTS) system outputs with human preferences has proven to be an effective approach for enhancing the robustness of language model-based TTS systems. Current approaches primarily focus on using preference data annotated at the utterance level. However, frequent issues that affect the listening experience often only arise in specific segments of audio samples, while other segments are well-generated. In this study, we propose a fine-grained preference optimization approach (FPO) to enhance the robustness of TTS systems. FPO focuses on addressing localized issues in generated samples rather than uniformly optimizing the entire utterance. Specifically, we first analyze the types of issues in generated samples, categorize them into two groups, and propose a selective training loss strategy to optimize preferences based on fine-grained labels for each issue type. Experimental results show that FPO enhances the robustness of zero-shot TTS systems by effectively addressing local issues, significantly reducing the bad case ratio, and improving intelligibility. Furthermore, FPO exhibits superior data efficiency compared with baseline systems, achieving similar performance with fewer training samples.

语音合成偏好优化零样本细粒度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。