arXiv:2607.19859cs.SD2026-07中稿 · ASRU 2025

用稀疏时间嵌入实现低延迟高鲁棒的语音合成,兼顾自然语调。

StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis

论文配图:StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis
图 1 · 摘自论文原文
  • 通过稀疏时间嵌入精细控制发音时长与语调,提升非自回归模型的灵活性。
  • 8300万参数模型实现实时因子0.08,延迟低于现有先进系统。
  • 适合移动端部署,兼顾音质、语调自然度和抗错能力,适合实际应用。

语音合成中鲁棒性、延迟与韵律之间的权衡仍是核心挑战。自回归模型虽保真度高,但速度慢且易出错;非自回归(NAR)模型虽快,却因固定对齐牺牲语调自然度。本文提出StellarTTS,一种基于稀疏时间嵌入策略的轻量化NAR语音合成框架,可精准调控音素时长、发音与韵律。同时引入语义感知编码器,支持单阶段高效解码。在稀疏时间嵌入条件下,83M参数的轻量级掩码生成式变换器实现0.08的实时因子(RTF)。实验表明,StellarTTS在延迟更低、鲁棒性更强的同时,保持了与当前最优系统相当的音频质量、韵律自然度和说话人相似性。

原文摘要 · Abstract (English)

The trade-off between robustness, latency, and prosody critically challenges text-to-speech (TTS) systems. Autoregressive models, despite fidelity, are slow and error-prone; non-autoregressive (NAR) alternatives, while fast, often sacrifice prosodic naturalness via rigid alignments. This paper introduces StellarTTS, a novel mobile-optimized NAR TTS framework based on a sparse temporal embedding strategy, enabling granular control of phoneme duration, pronunciation, and prosody. Furthermore, we propose a semantic-aware codec that facilitates efficient single-stage decoding. Conditioned on the sparse temporal embedding, our 83M-parameter lightweight masked generative transformer achieves a real-time factor (RTF) of 0.08. Experiments demonstrate that StellarTTS attains lower latency and stronger robustness compared to state-of-the-art TTS systems, while maintaining competitive performance in audio quality, prosodic naturalness, and speaker similarity.

语音合成非自回归低延迟稀疏嵌入

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。