arXiv:2604.19797eess.AScs.AI2026-04被引 1

用自适应置信度训练提升南印度语医学语音识别准确率

Enhancing ASR Performance in the Medical Domain for Dravidian Languages

  • 融合真实与合成语音,通过动态置信度加权训练
  • 泰卢固语WER降8.5%,卡纳达语WER降6.3%
  • 适合低资源、形态复杂的医疗语音场景

针对泰卢固语和卡纳达语等低资源南印度语言在医学领域的语音识别挑战,本文提出一种新型置信度感知训练框架。该框架结合静态感知与声学相似性指标及动态模型熵,实现对真实与合成语音数据的混合加权。相比直接微调,采用固定权重与可学习权重两种策略,在泰卢固语与卡纳达语医学数据集上进行训练。同时使用5-gram KenLM语言模型进行解码后修正。结果表明,采用可学习权重的混合置信度方法显著降低错误率:泰卢固语词错误率(WER)从24.3%降至15.8%(绝对下降8.5%),卡纳达语从31.7%降至25.4%(绝对下降6.3%),均显著优于标准微调基线。证实了自适应置信度训练与统计语言模型结合在形态复杂、领域专精的南印度语中具有优越性能。

原文摘要 · Abstract (English)

Automatic Speech Recognition (ASR) for low-resource Dravidian languages like Telugu and Kannada faces significant challenges in specialized medical domains due to limited annotated data and morphological complexity. This work proposes a novel confidence-aware training framework that integrates real and synthetic speech data through a hybrid confidence mechanism combining static perceptual and acoustic similarity metrics with dynamic model entropy. Unlike direct fine-tuning approaches, the proposed methodology employs both fixed-weight and learnable-weight confidence aggregation strategies to guide sample weighting during training, enabling effective utilization of heterogeneous data sources. The framework is evaluated on Telugu and Kannada medical datasets containing both real recordings and TTS-generated synthetic speech. A 5-gram KenLM language model is applied for post-decoding correction. Results show that the hybrid confidence-aware approach with learnable weights substantially reduces recognition errors: Telugu Word Error Rate (WER) decreases from 24.3% to 15.8% (8.5% absolute improvement), while Kannada WER drops from 31.7% to 25.4% (6.3% absolute improvement), both significantly outperforming standard fine-tuning baselines. These findings confirm that combining adaptive confidence-aware training with statistical language modeling delivers superior performance for domain-specific ASR in morphologically complex Dravidian languages.

语音识别医学ASR低资源语言置信度加权

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。