为卢旺达语和斯瓦希里语设计低延迟语音转写与合成系统
Edge-Based Speech Transcription and Synthesis for Kinyarwanda and Swahili Languages
- 采用边缘-云端协同推理,分担计算负载降低延迟
- 在1.7GHz CPU设备上,270字符处理时间小于1分钟
- 内存占用压缩超9%,适合资源受限地区使用
本文提出一种新型语音转写与合成框架,利用边缘-云并行计算提升卢旺达语和斯瓦希里语的处理速度与可及性。针对东非国家技术基础设施有限、语言处理工具稀缺的问题,框架采用Whisper和SpeechT5预训练模型实现语音转文本(STT)与文本转语音(TTS)。通过级联机制将模型推理任务分配至边缘设备与云端,有效降低延迟与资源消耗。实验表明,在1.7 GHz CPU与1 MB/s网络带宽下,边缘设备内存占用分别压缩9.5%(SpeechT5)和14%(Whisper),峰值仅149 MB;270字符文本的语音转写与语音合成可在60秒内完成。基于肯尼亚真实调研数据,该架构具备高准确率与良好响应速度,适合作为本地化语音服务的可靠平台。
原文摘要 · Abstract (English)
This paper presents a novel framework for speech transcription and synthesis, leveraging edge-cloud parallelism to enhance processing speed and accessibility for Kinyarwanda and Swahili speakers. It addresses the scarcity of powerful language processing tools for these widely spoken languages in East African countries with limited technological infrastructure. The framework utilizes the Whisper and SpeechT5 pre-trained models to enable speech-to-text (STT) and text-to-speech (TTS) translation. The architecture uses a cascading mechanism that distributes the model inference workload between the edge device and the cloud, thereby reducing latency and resource usage, benefiting both ends. On the edge device, our approach achieves a memory usage compression of 9.5% for the SpeechT5 model and 14% for the Whisper model, with a maximum memory usage of 149 MB. Experimental results indicate that on a 1.7 GHz CPU edge device with a 1 MB/s network bandwidth, the system can process a 270-character text in less than a minute for both speech-to-text and text-to-speech transcription. Using real-world survey data from Kenya, it is shown that the cascaded edge-cloud architecture proposed could easily serve as an excellent platform for STT and TTS transcription with good accuracy and response time.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。