arXiv:2606.09335eess.AS2026-06

分析印度语语音识别在不同条件下的表现差异,揭示关键影响因素。

Factors affecting ASR performance: A study using state of the art ASR models in Indic Languages

论文配图:Factors affecting ASR performance: A study using state of the art ASR models in Indic Languages
图 1 · 摘自论文原文
  • 用多个开源模型在零样本设置下测试多种印地语语音数据集。
  • 发现说话人特征与音频处理方式显著影响识别错误率。
  • 适合关注印度语语音识别落地的开发者和研究者参考。

语音识别性能在语言、说话人和录音条件下存在差异,但针对印地语族语言的系统性分析仍有限。本文对多个开源语音识别模型在多种印度语音数据集上的零样本测试结果进行大规模分析,涵盖印地语、孟加拉语、卡纳达语、泰卢固语和马拉地语。研究考察了语言学、说话人层级和声学因素的影响,分析了词错误率(WER)与说话人特征(如平均词长、语速、语句时长)之间的相关性。针对印地语,进一步分析了电话编码器、比特深度、重采样和背景噪声等音频因素。结果揭示了跨语言共性与语言特异性敏感性,表明说话人行为和信号处理选择会显著影响真实场景中印地语语音识别的鲁棒性。

原文摘要 · Abstract (English)

ASR performance varies across languages, speakers, and recording conditions, yet systematic analysis for Indic languages remain limited. We present a large-scale study of decoded outputs from multiple open-source ASR models evaluated on diverse Indian speech datasets in zero-shot settings. We analyze linguistic, speaker-level, and acoustic factors across Hindi, Bengali, Kannada, Telugu, and Marathi. We examine correlations between WER and speaker traits such as average word length, speaking rate, and utterance duration across multiple model dataset pairs. For Hindi, we further analyze audio factors including telephone codecs, bit depth, resampling, and background noise. Results reveal both cross lingual patterns and language-specific sensitivities, showing how speaker behavior and signal processing choices affect ASR robustness in real world Indic scenarios.

语音识别印地语多语言零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。