用稀疏时间嵌入实现低延迟高鲁棒的语音合成,兼顾自然语调。
StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis

- 通过稀疏时间嵌入精细控制发音时长与语调,提升非自回归模型的灵活性。
- 8300万参数模型实现实时因子0.08,延迟低于现有先进系统。
- 适合移动端部署,兼顾音质、语调自然度和抗错能力,适合实际应用。
语音合成中鲁棒性、延迟与韵律之间的权衡仍是核心挑战。自回归模型虽保真度高,但速度慢且易出错;非自回归(NAR)模型虽快,却因固定对齐牺牲语调自然度。本文提出StellarTTS,一种基于稀疏时间嵌入策略的轻量化NAR语音合成框架,可精准调控音素时长、发音与韵律。同时引入语义感知编码器,支持单阶段高效解码。在稀疏时间嵌入条件下,83M参数的轻量级掩码生成式变换器实现0.08的实时因子(RTF)。实验表明,StellarTTS在延迟更低、鲁棒性更强的同时,保持了与当前最优系统相当的音频质量、韵律自然度和说话人相似性。
原文摘要 · Abstract (English)
The trade-off between robustness, latency, and prosody critically challenges text-to-speech (TTS) systems. Autoregressive models, despite fidelity, are slow and error-prone; non-autoregressive (NAR) alternatives, while fast, often sacrifice prosodic naturalness via rigid alignments. This paper introduces StellarTTS, a novel mobile-optimized NAR TTS framework based on a sparse temporal embedding strategy, enabling granular control of phoneme duration, pronunciation, and prosody. Furthermore, we propose a semantic-aware codec that facilitates efficient single-stage decoding. Conditioned on the sparse temporal embedding, our 83M-parameter lightweight masked generative transformer achieves a real-time factor (RTF) of 0.08. Experiments demonstrate that StellarTTS attains lower latency and stronger robustness compared to state-of-the-art TTS systems, while maintaining competitive performance in audio quality, prosodic naturalness, and speaker similarity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。