arXiv:2602.03868eess.AScs.AI2026-02被引 1

为印度农业场景下的多语言语音识别建立基准,提升方言识别准确率。

Benchmarking Automatic Speech Recognition for Indian Languages in Agricultural Contexts

  • 构建跨三种语言的农业语音识别评估框架,引入领域加权词错误率等新指标。
  • 测试10,934段录音,发现印地语表现最佳(WER 16.2%),奥迪亚语最差(最低WER 35.1%)。
  • 提出说话人分离+优选策略,可降低多说话人音频的错误率最高达66%。

印度农业咨询数字化需要具备在多种印度语言中准确转录领域术语的稳健自动语音识别(ASR)系统。本文提出一个用于评估印地语、泰卢固语和奥迪亚语在农业场景下ASR性能的基准框架。我们引入农业加权词错误率(AWWER)和领域特定效用评分等评估指标,以补充传统指标。对10,934段音频录音(每段由最多10个ASR模型转录)的评估显示,不同语言和模型间性能存在差异:印地语整体表现最佳(词错误率WER: 16.2%),而奥迪亚语挑战最大(最优WER: 35.1%,仅在使用说话人分离时达成)。我们分析了真实农业田间录音中存在的音频质量挑战,并证明结合说话人分离与最佳说话人选择可显著降低多说话人录音的词错误率(最高可达66%,取决于多说话人音频比例)。我们识别出农业术语中的重复错误模式,并提供改进低资源农业领域ASR系统的实用建议。该研究为未来农业语音识别发展建立了基线基准。

原文摘要 · Abstract (English)

The digitization of agricultural advisory services in India requires robust Automatic Speech Recognition (ASR) systems capable of accurately transcribing domain-specific terminology in multiple Indian languages. This paper presents a benchmarking framework for evaluating ASR performance in agricultural contexts across Hindi, Telugu, and Odia languages. We introduce evaluation metrics including Agriculture Weighted Word Error Rate (AWWER) and domain-specific utility scoring to complement traditional metrics. Our evaluation of 10,934 audio recordings, each transcribed by up to 10 ASR models, reveals performance variations across languages and models, with Hindi achieving the best overall performance (WER: 16.2%) while Odia presents the greatest challenges (best WER: 35.1%, achieved only with speaker diarization). We characterize audio quality challenges inherent to real-world agricultural field recordings and demonstrate that speaker diarization with best-speaker selection can substantially reduce WER for multi-speaker recordings (upto 66% depending on the proportion of multi-speaker audio). We identify recurring error patterns in agricultural terminology and provide practical recommendations for improving ASR systems in low-resource agricultural domains. The study establishes baseline benchmarks for future agricultural ASR development.

语音识别农业AI多语言低资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。