arXiv:2507.02407cs.CLcs.LG2025-07中稿 · presentation at th…被引 4

对比7种Akan语音识别模型在不同语境下的表现,发现跨领域泛化能力差。

Benchmarking Akan ASR Models Across Domain-Specific Datasets: A Comparative Evaluation of Performance, Scalability, and Adaptability

  • 用4个不同领域的Akan语料测试Transformer架构模型
  • 跨领域时词错误率显著上升,最佳表现仅限训练域
  • Whisper更流畅但易误导,Wav2Vec2更准确但难解释,适合不同场景

现有自动语音识别(ASR)研究多基于同域数据集评估模型,却很少考察其在多样语音场景中的泛化能力。本研究通过基准测试7个基于Transformer架构的Akan ASR模型(如Whisper和Wav2Vec2),使用4个Akan语料库评估其性能,涵盖文化相关图像描述、非正式对话、圣经诵读及自发金融对话等不同领域。结果表明,模型在训练域内表现最优,跨域时词错误率(WER)和字符错误率(CER)显著上升,存在明显领域依赖性。Whisper微调模型生成更流畅但可能误导的转录错误,而Wav2Vec2在陌生输入下产生更明显但难以解读的输出。这一可读性与透明性之间的权衡需在低资源语言(LRL)应用中考虑。研究强调需发展针对性领域适配技术、自适应路由策略及多语言训练框架以提升Akan及其他LRL的语音识别能力。

原文摘要 · Abstract (English)

Most existing automatic speech recognition (ASR) research evaluate models using in-domain datasets. However, they seldom evaluate how they generalize across diverse speech contexts. This study addresses this gap by benchmarking seven Akan ASR models built on transformer architectures, such as Whisper and Wav2Vec2, using four Akan speech corpora to determine their performance. These datasets encompass various domains, including culturally relevant image descriptions, informal conversations, biblical scripture readings, and spontaneous financial dialogues. A comparison of the word error rate and character error rate highlighted domain dependency, with models performing optimally only within their training domains while showing marked accuracy degradation in mismatched scenarios. This study also identified distinct error behaviors between the Whisper and Wav2Vec2 architectures. Whereas fine-tuned Whisper Akan models led to more fluent but potentially misleading transcription errors, Wav2Vec2 produced more obvious yet less interpretable outputs when encountering unfamiliar inputs. This trade-off between readability and transparency in ASR errors should be considered when selecting architectures for low-resource language (LRL) applications. These findings highlight the need for targeted domain adaptation techniques, adaptive routing strategies, and multilingual training frameworks for Akan and other LRLs.

语音识别低资源语言模型泛化领域适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。