arXiv:2606.20266eess.AS2026-06中稿 · Interspeech 2026被引 1

无需参考文本,用语音特征实现更自然的零样本语音合成

Transcript-Free Flow-Matching Text-to-Speech via Speech Feature Conditioning

论文配图:Transcript-Free Flow-Matching Text-to-Speech via Speech Feature Conditioning
图 1 · 摘自论文原文
  • 用轻量适配器将自监督语音表征映射到文本条件空间
  • 对发音障碍者合成错误率从24.6%降至10.4%,优于真实文本基线
  • 不依赖参考文本,适合口音或言语障碍者场景

近期基于流匹配的文本到语音模型(如F5-TTS)在推理时依赖外部语音识别系统获取参考文本,导致对有口音或言语障碍者的零样本语音合成性能脆弱,而这正是最需要该技术的场景。此外,我们发现基于文本的参考条件会将异常语音模式传播至合成结果,即使有真实文本也难以避免。为此,我们提出RTFree-F5,用连续的自监督语音表示替代参考文本,并通过轻量适配器将其映射到F5-TTS的文本条件空间,同时复用预训练模型权重。在发音障碍语音上,RTFree-F5将词错误率(WER)从24.6%降至10.4%,超过使用真实文本的基线,同时提升自然度,并在标准基准上保持竞争力,且无需任何参考文本。

原文摘要 · Abstract (English)

Recent flow-matching text-to-speech (TTS) models, such as F5-TTS, rely on a reference transcript at inference time, obtained from an external ASR system. This dependency makes zero-shot TTS brittle for accented or dysarthric speakers, precisely the scenarios where it is most needed. Moreover, we find that text-based reference conditioning can propagate atypical acoustic patterns from atypical speech into synthesis, even when ground-truth transcripts are available. To address this, we propose RTFree-F5, which replaces the reference transcript with continuous self-supervised speech representations mapped into F5-TTS's text-conditioning space via a lightweight adapter, while reusing the pretrained checkpoint. On dysarthric speech, RTFree-F5 reduces WER from 24.6% to 10.4%, surpassing even the ground-truth reference transcript baselines, while improving naturalness and remaining competitive on standard benchmarks without requiring any reference transcript.

语音合成零样本自监督口音适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。