arXiv:2409.15759cs.SDeess.AS2024-09被引 3

用自动引导提升低资源语音适配效果,让小模型也能搞定陌生声音。

VoiceGuider: Enhancing Out-of-Domain Performance in Parameter-Efficient Speaker-Adaptive Text-to-Speech via Autoguidance

  • 通过自动引导机制增强LoRA微调的语音适应能力
  • 在极端域外语音上表现显著优于普通参数高效方法
  • 适合资源受限场景下需跨说话人适配的语音合成应用

在参数高效的LoRA微调应用于说话人自适应文本到语音模型时,对域外说话人的适配性能可能低于全量微调模型。本文提出VoiceGuider,一种基于自动引导的参数高效语音合成系统,旨在缩小与全量微调模型的性能差距。我们系统探索了多种强化自动引导的策略,最终确定最优方案。实验表明,该方法在极端域外语音数据上表现出稳健的适应能力。演示页面提供可听样本。

原文摘要 · Abstract (English)

When applying parameter-efficient finetuning via LoRA onto speaker adaptive text-to-speech models, adaptation performance may decline compared to full-finetuned counterparts, especially for out-of-domain speakers. Here, we propose VoiceGuider, a parameter-efficient speaker adaptive text-to-speech system reinforced with autoguidance to enhance the speaker adaptation performance, reducing the gap against full-finetuned models. We carefully explore various ways of strengthening autoguidance, ultimately finding the optimal strategy. VoiceGuider as a result shows robust adaptation performance especially on extreme out-of-domain speech data. We provide audible samples in our demo page.

语音合成参数高效自适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。