arXiv:2502.07562cs.SDcs.AI2025-02被引 3

用低秩适配让语音合成更精准模仿真实环境下的说话声

LoRP-TTS: Low-Rank Personalized Text-To-Speech

  • 用低秩适配技术,仅需一段嘈杂环境录音即可生成逼真语音
  • 在非专业录音条件下,说话人相似度提升达30个百分点
  • 适合需要个性化语音生成的智能客服、虚拟助手场景

语音合成模型将文本转化为自然语音。早期模型仅支持单一说话人,近年发展出零样本系统,能通过额外语音提示生成多样说话人声音。然而,这些系统在处理与训练数据差异大的非专业录音时仍表现不佳。本文证明,采用低秩适配(LoRA)技术可成功使用单段嘈杂环境中的自发语音作为提示。该方法使说话人相似度最高提升30个百分点,同时保持内容准确性和语音自然度。这一进展对构建真正多样化的语音语料库具有重要意义,对所有语音相关任务都至关重要。

原文摘要 · Abstract (English)

Speech synthesis models convert written text into natural-sounding audio. While earlier models were limited to a single speaker, recent advancements have led to the development of zero-shot systems that generate realistic speech from a wide range of speakers using their voices as additional prompts. However, they still struggle with imitating non-studio-quality samples that differ significantly from the training datasets. In this work, we demonstrate that utilizing Low-Rank Adaptation (LoRA) allows us to successfully use even single recordings of spontaneous speech in noisy environments as prompts. This approach enhances speaker similarity by up to $30pp$ while preserving content and naturalness. It represents a significant step toward creating truly diverse speech corpora, that is crucial in all speech-related tasks.

语音合成低秩适配个性化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。