arXiv:2510.05799cs.CLcs.AI2025-10ACL

无需配对数据,直接优化语音合成中每个音素的自然度。

Data-efficient Targeted Token-level Preference Optimization for LLM-based Text-to-Speech

  • 不依赖成对样本,直接在音素粒度上进行偏好优化。
  • 日语语音合成错误率降低54%,发音准确率提升39%。
  • 自动为关键音素赋予12.8倍更强奖励,适合高精度语音生成任务。

通过偏好优化对齐文本到语音(TTS)系统输出与人类反馈,已被证明能有效提升基于语言模型的TTS模型的鲁棒性和自然度。现有方法主要依赖于话语级别的理想与非理想样本配对,但此类配对在语音输出数据中往往稀缺,且话语级建模无法实现精确发音对齐所需的细粒度音素级优化。本研究提出TKTO方法,消除对配对数据的需求,实现更高效的数据利用,并直接针对音素级别单位进行优化,无需音素级标注即可自动提供细粒度对齐信号。实验表明,TKTO将挑战性的日语语音合成准确率提高39%,字符错误率(CER)降低54%,并自动为目标音素赋予12.8倍更强的奖励信号。

原文摘要 · Abstract (English)

Aligning text-to-speech (TTS) system outputs with human feedback through preference optimization has been shown to effectively improve the robustness and naturalness of language model-based TTS models. Current approaches primarily require paired desirable and undesirable samples at the utterance level. However, such pairs are often limited in TTS output data, and utterance-level formulation prevents fine-grained token-level optimization needed for accurate pronunciation alignment. In this study, we propose TKTO that eliminates the need for paired data, enabling a more data-efficient training paradigm, and directly targets token-level units, automatically providing fine-grained alignment signals without token-level annotations. TKTO improves the challenging Japanese TTS accuracy by 39% and reduces CER by 54%, automatically assigning 12.8 times stronger reward to targeted tokens.

语音合成偏好优化高效训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。