arXiv:2507.01348eess.AScs.SD2025-07被引 2

用大模型统一解决外语口音转换与语音合成问题

SpeechAccentLLM: A Unified Framework for Foreign Accent Conversion and Text to Speech

  • 提出SpeechCodeVAE,用CTC实现语音内容分词并保持时序连贯性
  • 多任务联合训练提升口音转换效果,加速收敛且音质更优
  • 引入SpeechRestorer修复大模型输出错误,增强语音韵律自然度

外语口音转换(FAC)在语音处理中仍具挑战。基于大语言模型(LLM)在文本转语音(TTS)任务中的成功,本研究探索将基于LLM的技术应用于FAC,提出SpeechAccentLLM框架。核心是SpeechCodeVAE,首个将连接时序分类(CTC)直接融入码本离散化以实现语音内容分词的模型。该架构生成具有独特“局部性”特性的编码,实验验证其在内容忠实度、时序连贯性和结构可恢复性之间达到最优权衡。为缓解FAC模块的数据稀缺问题,采用多任务学习策略,联合训练FAC与TTS模块。该方法不仅缓解数据限制,还实现更快收敛和更优语音质量。此外,基于离散语音表示的显著特性,提出SpeechRestorer后处理架构,有效缓解LLM推理中常见的随机误差,提升语调连贯性,经消融实验验证有效。

原文摘要 · Abstract (English)

Foreign accent conversion (FAC) in speech processing remains a challenging task. Building on the remarkable success of large language models (LLMs) in Text-to-Speech (TTS) tasks, this study investigates the adaptation of LLM-based techniques for FAC, which we term SpeechAccentLLM. At the core of this framework, we introduce SpeechCodeVAE, the first model to integrate connectionist temporal classification (CTC) directly into codebook discretization for speech content tokenization. This novel architecture generates tokens with a unique "locality" property, as validated by experiments demonstrating optimal trade-offs among content faithfulness, temporal coherence, and structural recoverability. Then, to address data scarcity for the FAC module, we adopted a multitask learning strategy that jointly trains the FAC and TTS modules. Beyond mitigating data limitations, this approach yielded accelerated convergence and superior speech quality compared to standalone FAC training. Moreover, leveraging the salient properties of our discrete speech representations, we introduce SpeechRestorer, a postprocessing architecture designed to refine LLM-generated outputs. This module effectively mitigates stochastic errors prevalent in LLM inference pipelines while enhancing prosodic continuity, as validated by ablation experiments.

语音合成口音转换大模型多任务学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。