arXiv:2512.00937eess.AS2025-12被引 2

解决阿拉伯语语音合成资源少、音韵复杂难题,提升语音自然度。

Arabic TTS with FastPitch: Reproducible Baselines, Adversarial Training, and Oversmoothing Analysis

  • 基于FastPitch架构构建可复现的阿拉伯语语音合成基线。
  • 提出倒谱域指标,揭示频谱预测中过度平滑问题及其训练过程影响。
  • 引入轻量对抗损失,显著减少语音平滑现象,适合多说话人合成研究者使用。

阿拉伯语文本转语音(TTS)因资源有限和复杂的音韵模式而面临挑战。本文基于FastPitch架构构建了可复现的阿拉伯语TTS基线,并提出倒谱域指标,用于分析梅尔频谱预测中的过度平滑现象。传统Lp重构损失虽生成平滑输出,但导致信息过平均;所提指标揭示其在训练过程中对时序与频谱的影响。为解决此问题,引入轻量级对抗频谱损失,训练稳定且显著降低过度平滑。此外,通过XTTSv2生成合成语音增强多说话人阿拉伯语TTS,提升语调多样性而不牺牲稳定性。代码、预训练模型及训练方案已公开:https://github.com/nipponjo/tts-arabic-pytorch。

原文摘要 · Abstract (English)

Arabic text-to-speech (TTS) remains challenging due to limited resources and complex phonological patterns. We present reproducible baselines for Arabic TTS built on the FastPitch architecture and introduce cepstral-domain metrics for analyzing oversmoothing in mel-spectrogram prediction. While traditional Lp reconstruction losses yield smooth but over-averaged outputs, the proposed metrics reveal their temporal and spectral effects throughout training. To address this, we incorporate a lightweight adversarial spectrogram loss, which trains stably and substantially reduces oversmoothing. We further explore multi-speaker Arabic TTS by augmenting FastPitch with synthetic voices generated using XTTSv2, resulting in improved prosodic diversity without loss of stability. The code, pretrained models, and training recipes are publicly available at: https://github.com/nipponjo/tts-arabic-pytorch.

语音合成阿拉伯语对抗训练FastPitch

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。