arXiv:2603.10904cs.SDcs.AI2026-03被引 1

LoRA微调让小模型语音克隆更逼真,数据越多样效果越明显。

When Fine-Tuning Fails and when it Generalises: Role of Data Diversity and Mixed Training in LLM-based TTS

  • 用LoRA微调语言模型,提升语音一致性与音质。
  • 语音质量提升最高达0.42分DNS-MOS,信噪比提高34%。
  • 适合追求低延迟、高音质的轻量化语音克隆场景。

大语言模型被越来越多地用作神经文本转语音系统的语义主干。然而,冻结的LLM表示无法充分建模说话人特有的声学和感知特征。我们对TTS系统中的语言模型主干进行微调实验,结果显示在语音克隆任务中能显著提升语音一致性和信噪比(SNR)。在多个说话人上,LoRA微调在三个互补维度均优于未微调的基线Qwen-0.5B模型:第一,感知质量显著提升,对于训练数据具有足够声学变异性的话者,DNS-MOS最高提升0.42分;第二,说话人保真度普遍提高,语音相似性持续增强,表明LoRA有效适应了说话人身份表征且未损害语言建模能力;第三,信号级质量在多数情况下改善,信噪比最高提升34%。关键发现是这些改进强烈依赖于训练数据特性:声学能量和感知质量变化大的话者,在DNS-MOS、语音相似性和信噪比上均获得同步提升。总体而言,本研究证明LoRA微调不仅是参数高效的优化方法,更是紧凑型LLM-TTS系统中实现更好说话人适配的有效机制。在具备足够多样性的训练数据支持下,经LoRA适配的Qwen-0.5B模型在感知质量、说话人相似性方面持续超越其冻结基线模型,且在量化后的GGUF模型下实现低延迟部署。

原文摘要 · Abstract (English)

Large language models are increasingly adopted as semantic backbones for neural text-to-speech systems. However, frozen LLM representations are insufficient for modeling speaker specific acoustic and perceptual characteristics. Our experiments involving fine tuning of the Language Model backbone of TTS show promise in improving the voice consistency and Signal to Noise ratio SNR in voice cloning task. Across multiple speakers LoRA finetuning consistently outperforms the non-finetuned base Qwen-0.5B model across three complementary dimensions of speech quality. First, perceptual quality improves significantly with DNS-MOS gains of up to 0.42 points for speakers whose training data exhibits sufficient acoustic variability. Second, speaker fidelity improves for all evaluated speakers with consistent increases in voice similarity indicating that LoRA effectively adapts speaker identity representations without degrading linguistic modeling. Third, signal level quality improves in most cases with signal to noise ratio increasing by as much as 34 percent. Crucially these improvements are strongly governed by the characteristics of the training data. Speakers with high variability in acoustic energy and perceptual quality achieve simultaneous gains in DNS-MOS voice similarity and SNR. Overall this work establishes that LoRA finetuning is not merely a parameter efficient optimization technique but an effective mechanism for better speaker level adaptation in compact LLM-based TTS systems. When supported by sufficiently diverse training data LoRA adapted Qwen-0.5B consistently surpasses its frozen base model in perceptual quality speaker similarity with low latency using GGUF model hosted in quantized form.

语音克隆LoRA微调小模型音质提升

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。