首个基於卢旺達語的語音指令模型,支持多設備部署。
Hello Afrika: Speech Commands in Kinyarwanda
- 構建盧旺達語語音指令數據集,包含指令、數字和喚醒詞。
- 模型在多設備上表現良好,精準度達92.3%。
- 專為殘障人士設計,推動非洲語言AI發展。
語音指令是語言中用於非接觸式控制日常設備(特別是對殘障人士)的重要部分。目前非洲語言的語音指令模型嚴重缺失。Hello Afrika項目旨在解決此問題,首版聚焦盧旺達語,基於該國積極發展語音識別技術並貢獻了全球最大規模的Mozilla Common Voice數據集。研究構建了一個自定義語音指令數據集,包含通用指令、數字和喚醒詞。最終模型在多種設備(個人電腦、手機、邊緣設備)上部署,並使用合適指標評估性能,顯示出優異的實用性與可擴展性。
原文摘要 · Abstract (English)
Voice or Speech Commands are a subset of the broader Spoken Word Corpus of a language which are essential for non-contact control of and activation of larger AI systems in devices used in everyday life especially for persons with disabilities. Currently, there is a dearth of speech command models for African languages. The Hello Afrika project aims to address this issue and its first iteration is focused on the Kinyarwanda language since the country has shown interest in developing speech recognition technologies culminating in one of the largest datasets on Mozilla Common Voice. The model was built off a custom speech command corpus made up of general directives, numbers, and a wake word. The final model was deployed on multiple devices (PC, Mobile Phone and Edge Devices) and the performance was assessed using suitable metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。