arXiv:2510.18938eess.AScs.AI2025-10被引 2

首个直接将口吃语音转为流畅语音的端到端模型,兼顾转录与修复。

StutterZero and StutterFormer: End-to-End Speech Conversion for Stuttering Transcription and Correction

  • 端到端波形到波形转换,联合优化转录与语音修复。
  • 相比Whisper-Medium,WER降低24%~28%,语义相似度提升31%~34%。
  • 适合语音治疗、无障碍交互等需要精准口吃处理的场景。

全球超7000万人有口吃问题,但现有自动语音系统常误读或无法准确转录不流畅语音。传统方法依赖手工特征或分阶段的语音识别(ASR)与文本转语音(TTS)流程,分离转录与音频重建,易放大失真。本文提出StutterZero与StutterFormer,首个端到端波形到波形模型,可直接将口吃语音转为流畅语音并同步生成转录文本。StutterZero采用卷积-双向LSTM编码器-解码器结构,StutterFormer则结合双流Transformer与共享声学-语言表示。两者在SEP-28K与LibriStutter合成数据上训练,并在FluencyBank未见说话人上评估。所有基准测试中,StutterZero相较领先模型Whisper-Medium实现24%的词错误率(WER)下降和31%的语义相似度提升(BERTScore);StutterFormer表现更优,达到28%的WER下降和34%的BERTScore提升。结果验证了直接端到端口吃转流畅语音的可行性,为包容性人机交互、语音治疗及无障碍AI系统带来新机遇。

原文摘要 · Abstract (English)

Over 70 million people worldwide experience stuttering, yet most automatic speech systems misinterpret disfluent utterances or fail to transcribe them accurately. Existing methods for stutter correction rely on handcrafted feature extraction or multi-stage automatic speech recognition (ASR) and text-to-speech (TTS) pipelines, which separate transcription from audio reconstruction and often amplify distortions. This work introduces StutterZero and StutterFormer, the first end-to-end waveform-to-waveform models that directly convert stuttered speech into fluent speech while jointly predicting its transcription. StutterZero employs a convolutional-bidirectional LSTM encoder-decoder with attention, whereas StutterFormer integrates a dual-stream Transformer with shared acoustic-linguistic representations. Both architectures are trained on paired stuttered-fluent data synthesized from the SEP-28K and LibriStutter corpora and evaluated on unseen speakers from the FluencyBank dataset. Across all benchmarks, StutterZero had a 24% decrease in Word Error Rate (WER) and a 31% improvement in semantic similarity (BERTScore) compared to the leading Whisper-Medium model. StutterFormer achieved better results, with a 28% decrease in WER and a 34% improvement in BERTScore. The results validate the feasibility of direct end-to-end stutter-to-fluent speech conversion, offering new opportunities for inclusive human-computer interaction, speech therapy, and accessibility-oriented AI systems.

语音修复端到端口吃处理语音合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。