arXiv:2509.02523cs.CLcs.LG2025-09被引 2

小模型专精少数语言,语音识别效果远超大模型。

Flavors of Moonshine: Tiny Specialized ASR Models for Edge Devices

  • 用真实、伪标签和合成数据训练单语小模型。
  • 误差率比Whisper Tiny低48%,超越9倍大的Whisper Small。
  • 适合资源少的语言,支持设备端实时语音识别。

我们提出Flavors of Moonshine,一套针对欠代表语言的微型自动语音识别(ASR)模型。主流观点认为多语言模型因跨语言语音相似性优于单语模型。我们挑战这一假设,发现对于参数量仅为2700万的小模型,通过精心平衡的高质量人工标注、伪标注和合成数据训练单语系统,可实现显著更优性能。平均而言,我们的模型误差率比同规模的Whisper Tiny低48%,超越9倍大的Whisper Small模型,多数情况下与28倍大的Whisper Medium模型持平或更优。这些成果推动了该尺寸模型的最先进水平,使此前支持不足的语言也能实现精准的设备端语音识别。我们已将阿拉伯语、中文、日语、韩语、乌克兰语和越南语的Moonshine模型以宽松开源许可发布。

原文摘要 · Abstract (English)

We present the Flavors of Moonshine, a suite of tiny automatic speech recognition (ASR) models specialized for a range of underrepresented languages. Prevailing wisdom suggests that multilingual ASR models outperform monolingual counterparts by exploiting cross-lingual phonetic similarities. We challenge this assumption, showing that for sufficiently small models (27M parameters), training monolingual systems on a carefully balanced mix of high-quality human-labeled, pseudo-labeled, and synthetic data yields substantially superior performance. On average, our models achieve error rates 48% lower than the comparably sized Whisper Tiny model, outperform the 9x larger Whisper Small model, and in most cases match or outperform the 28x larger Whisper Medium model. These results advance the state of the art for models of this size, enabling accurate on-device ASR for languages that previously had limited support. We release Arabic, Chinese, Japanese, Korean, Ukrainian, and Vietnamese Moonshine models under a permissive open-source license.

语音识别边缘计算小模型多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。