用语音合成生成多语言指令数据,解决边缘设备关键词识别缺数据难题
SynTTS-Commands: A Public Dataset for On-Device KWS via TTS-Synthesized Multilingual Speech
- 全用TTS合成英语和中文语音指令,突破真人录制瓶颈
- 在多种轻量模型上实现99.5%(英)和98%(中)识别准确率
- 适合做低功耗、本地化语音交互的开发者使用
面向超低功耗硬件的高性能本地化关键词识别(KWS)系统发展受限于专用多指令训练数据稀缺。传统真人录音成本高、速度慢且难以扩展。本文提出SynTTS-COMMANDS,一个完全通过先进文本转语音(TTS)技术生成的多语言语音指令数据集。利用CosyVoice 2模型与公开语料库中的说话人嵌入,构建了英语和中文指令的可扩展集合。在一系列高效声学模型上的广泛基准测试表明,该合成数据集使英语命令识别准确率最高达99.5%,中文达98%。结果有力验证了合成语音可有效替代真人录音用于训练KWS分类器。本工作直接缓解了TinyML领域的数据瓶颈,为资源受限边缘设备上构建私密、低延迟、低功耗语音界面提供了可行、可扩展的基础。数据集与源码已公开于https://github.com/lugan113/SynTTS-Commands-Official。
原文摘要 · Abstract (English)
The development of high-performance, on-device keyword spotting (KWS) systems for ultra-low-power hardware is critically constrained by the scarcity of specialized, multi-command training datasets. Traditional data collection through human recording is costly, slow, and lacks scalability. This paper introduces SYNTTS-COMMANDS, a novel, multilingual voice command dataset entirely generated using state-of-the-art Text-to-Speech (TTS) synthesis. By leveraging the CosyVoice 2 model and speaker embeddings from public corpora, we created a scalable collection of English and Chinese commands. Extensive benchmarking across a range of efficient acoustic models demonstrates that our synthetic dataset enables exceptional accuracy, achieving up to 99.5\% on English and 98\% on Chinese command recognition. These results robustly validate that synthetic speech can effectively replace human-recorded audio for training KWS classifiers. Our work directly addresses the data bottleneck in TinyML, providing a practical, scalable foundation for building private, low-latency, and energy-efficient voice interfaces on resource-constrained edge devices. The dataset and source code are publicly available at https://github.com/lugan113/SynTTS-Commands-Official.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。