让高音质语音在低功耗设备上实时运行,解决无障碍应用的延迟与体积难题。
Compact Neural TTS Voices for Accessibility
- 设计轻量级神经语音合成模型,兼顾音质与效率。
- 实现15毫秒级延迟,支持多语音预装于低配设备。
- 适合移动无障碍应用、离线语音助手等场景使用。
当前无障碍应用中的文本转语音技术主要分为两类:(i) 基于设备的统计参数语音合成(SPSS)或单元选择(USEL),具有低延迟和小存储占用,但音质自然度较差;(ii) 基于云端的神经语音合成(Neural TTS),音质优异,但延迟高、响应慢,难以用于实际场景。近期虽有神经TTS模型可部署于手持设备,但仍存在延迟高于SPSS/USEL、存储占用大而无法同时预装多个语音的问题。本文提出一种高质量紧凑型神经TTS系统,在低功耗设备上实现约15毫秒的延迟,同时保持极低的磁盘占用,可支持多语音预安装,适用于实际部署的无障碍应用。
原文摘要 · Abstract (English)
Contemporary text-to-speech solutions for accessibility applications can typically be classified into two categories: (i) device-based statistical parametric speech synthesis (SPSS) or unit selection (USEL) and (ii) cloud-based neural TTS. SPSS and USEL offer low latency and low disk footprint at the expense of naturalness and audio quality. Cloud-based neural TTS systems provide significantly better audio quality and naturalness but regress in terms of latency and responsiveness, rendering these impractical for real-world applications. More recently, neural TTS models were made deployable to run on handheld devices. Nevertheless, latency remains higher than SPSS and USEL, while disk footprint prohibits pre-installation for multiple voices at once. In this work, we describe a high-quality compact neural TTS system achieving latency on the order of 15 ms with low disk footprint. The proposed solution is capable of running on low-power devices.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。