Sylber 2.0用音节实现高效通用语音建模,支持多语言高保真还原。
Sylber 2.0: A Universal Syllable Embedding
- 自监督学习音节级语音编码,实现约5Hz低频压缩。
- 在多语言与多种语调下保持语音与语义细节,性能媲美高频基线模型。
- 适用于低资源语音识别与轻量级语音合成,适合多语言应用者。
构建语音建模需要高效且通用的语音标记。近期研究提出音节作为低时间分辨率下的有前景语音标记,但现有模型仅限于英语,难以捕捉充分的声学细节。为填补这一空白,我们提出 Sylber 2.0,一种自监督音节级语音编码框架,可实现高效的时序压缩与高保真重建。该模型在多语言和多种表达风格中均实现约5 Hz的极低标记频率,同时保留语言与声学细节。实验表明,其性能与基于高频基线的模型相当。此外,Sylber 2.0 支持高效语音合成,仅用7200万参数即可生成与当前最优模型相当的语音清晰度与质量。更重要的是,其通用性在低资源语音识别中优于以往语音编码框架。综上,我们建立了一种适用于通用口语语言的有效音节级抽象表示。
原文摘要 · Abstract (English)
Scaling spoken language modeling requires speech tokens that are both efficient and universal. Recent work has proposed syllables as promising speech tokens at low temporal resolution, but existing models are constrained to English and fail to capture sufficient acoustic detail. To address this gap, we present Sylber 2.0, a self-supervised framework for coding speech at the syllable level that enables efficient temporal compression and high-fidelity reconstruction. Sylber 2.0 achieves a very low token frequency around 5 Hz, while retaining both linguistic and acoustic detail across multiple languages and expressive styles. Experiments show that it performs on par with previous models operating on high-frequency baselines. Furthermore, Sylber 2.0 enables efficient TTS modeling which can generate speech with competitive intelligibility and quality with SOTA models using only 72M parameters. Moreover, the universality of Sylber 2.0 provides more effective features for low resource ASR than previous speech coding frameworks. In sum, we establish an effective syllable-level abstraction for general spoken language.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。