arXiv:2607.18912cs.CL2026-07

用多语言语音模型适配肯尼亚三种本土语言,解决数据不一致问题。

From a Multilingual Streaming ASR Backbone to Kenyan-Language Systems: Data-Centric Adaptation of Nemotron 3.5 for Kikuyu, Dholuo, and Kalenjin

  • 基于预训练模型全参数微调,保留流式推理能力
  • 基库尤语和道鲁语分别达到42.97%和33.98%的词错误率
  • 提供可审计的数据处理流程,适合低资源语言研究者

非洲语言自动语音识别受限于拼写不一致、标注噪声、音频缺失、说话人与领域分布失衡,以及评估方式与部署场景不符。本文开展端到端工程研究,将NVIDIA Nemotron 3.5 Streaming 0.6B ASR模型适配基库尤语、道鲁语和卡伦金语。从已适配肯尼亚斯瓦希里语的检查点出发,保留其缓存感知的FastConformer RNN-T、提示条件化与流式解码器,进行全参数微调。涵盖语料审核、Unicode归一化、分割检查、时长过滤、低速继续训练、基于验证的检查点选择、真实流式评估、噪声保留与独立服务等流程。在内部自适应评估集上(上下文[56,13]),选定的基库尤语和道鲁语模型分别取得42.97%与33.98%的词错误率(WER);道鲁语在冻结历史标签策略下,实现9.59%字符错误率(CER)与8.13%无空格字符错误率(no-space CER);基库尤语为7.79%无空格CER。卡伦金语仍处于进展中:在2,411条清洁诊断子集上(排除长停顿标注、含数字参考及短于三词目标),v1-v版本达68.74% WER。其检查点选择使用混合源验证清单(含测试来源样本),故该分数非独立泛化估计。此外报告了非语音标签、短语音过生成、边界敏感的WER及云作业生命周期失败等负面发现。因内部评估集重复使用、自适应咨询与归一化方式不同于公开基准,未声称达到最先进水平。本工作提供在不放弃流式约束前提下,将多语言模型转化为语言专用系统的可审计路径。

原文摘要 · Abstract (English)

Automatic speech recognition (ASR) for African languages is constrained by orthographic inconsistency, annotation artifacts, missing audio, speaker and domain imbalance, and evaluation procedures that differ from deployment. We present an end-to-end engineering study adapting NVIDIA Nemotron 3.5 ASR Streaming 0.6B to Kikuyu, Dholuo, and Kalenjin. Starting from a Kenyan Swahili-adapted checkpoint, we retain its cache-aware FastConformer RNN-T, prompt conditioning, and streaming decoder during full-parameter fine-tuning. The study covers corpus auditing, Unicode normalization, split checks, duration filtering, low-rate continuation, validation-based checkpoint selection, true-streaming evaluation, artifact preservation, and isolated serving. On internal, adaptively consulted evaluation sets excluded from gradient updates at context [56,13], selected Kikuyu and Dholuo models achieve 42.97% and 33.98% WER, respectively. Dholuo records 9.59% CER and 8.13% no-space CER under its frozen historical label policy; Kikuyu records 7.79% no-space CER. Kalenjin remains a work in progress: v1-v reaches 68.74% WER on a 2,411-row clean-v3 diagnostic subset excluding long-pause annotations, digit-bearing references, and targets shorter than three tokens. Its checkpoint selection used a mixed-source validation manifest containing test-origin rows, so the score is not an independent generalization estimate. We also report negative findings involving non-speech labels, short-utterance over-generation, boundary-sensitive WER, and cloud job-lifecycle failures. We make no state-of-the-art claim because the internal sets, repeated consultation, and normalization differ from public benchmarks. This work provides an auditable account of adapting a multilingual streaming model into language-specific systems without discarding streaming constraints.

语音识别非洲语言流式模型数据适配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。