arXiv:2501.14273eess.AScs.SD2025-01

针对大模型语音合成的语气与音色适配难题,提出高效微调方法。

Efficient Emotion and Speaker Adaptation in LLM-Based TTS via Characteristic-Specific Partial Fine-Tuning

  • 按特征重要性动态选择两层进行微调,仅更新约8%参数
  • 在11个数据集上达到全量微调的音色/语义保真度
  • 训练速度提升2倍,显著缓解灾难性遗忘问题

基于大语言模型的语音合成系统虽具备零样本情感与说话人克隆能力,但在未见领域中克隆保真度和发音清晰度会下降。微调对适应至关重要,但传统统一微调方式忽略参数贡献差异。有限数据下的统一微调导致训练缓慢且引发灾难性遗忘,影响发音准确性。为此,本文提出特征特定的部分微调策略CSP-FT。通过加权求和动态分析各层贡献,仅选择捕捉最多和最少情感与说话人信息的两层进行微调,最大化前者的利用效率,同时显式强化后者的能力。在包含11个数据集的联合语料库上的实验表明,CSP-FT在仅更新约8%参数的情况下,性能匹配或超越全量微调,在保持音色与语义保真度的同时,训练速度提升约2倍,并显著缓解灾难性遗忘。

原文摘要 · Abstract (English)

While LLM-based TTS models exhibit zero-shot emotion and speaker cloning, their cloning fidelity and pronunciation clarity degrade on unseen domains. Fine-tuning is essential for adaptation, yet uniform approaches overlook specific parameter contributions. Uniform tuning on limited data causes slow training and catastrophic forgetting, leading to degraded pronunciation accuracy. To address this, we propose CSP-FT, a characteristic-specific partial fine-tuning strategy. By dynamically analyzing layer contributions via a weighted sum, we selectively fine-tune only the two layers capturing the most and least emotion and speaker information, maximizing the utility of the former while explicitly strengthening the capacity of the latter. Experiments on a combined corpus of 11 datasets show CSP-FT matches or exceeds the fidelity and intelligibility of full fine-tuning while updating only ~8% of parameters, accelerating training by ~2x, and significantly mitigating catastrophic forgetting.

语音合成微调优化大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。