arXiv:2501.08616eess.AScs.SD2025-01

在极低资源下实现14种非洲语言识别,仅用开发集数据

IITKGP-ABSP Submission to LRE22: Language Recognition in Low-Resource Settings

  • 仅用开发集数据,不依赖预训练模型,通过音频增强与分类器融合
  • 在开发集上达到EER 11.43%、成本指标0.41的高效表现
  • 适合计算资源或存储受限场景,轻量级部署友好

本文详细描述了IITKGP-ABSP实验室在NIST语言识别评估(LRE22)中的系统方案。目标是识别14种低资源非洲语言。尽管NIST提供了额外训练和开发数据,但本研究在极端低资源约束下开展:主提交方案仅使用LRE22开发集中的14种目标语言语句,且禁止使用任何预训练模型进行特征提取或分类器微调。为应对低资源挑战,系统采用多样化的音频增强策略,并结合分类器融合机制。在满足所有约束条件下,该方法在开发集上实现了EER 11.43%和成本指标0.41的表现。对于计算资源有限或网络存储受限的用户,该系统可提供高效的语音语言识别性能。

原文摘要 · Abstract (English)

This is the detailed system description of the IITKGP-ABSP lab's submission to the NIST language recognition evaluation (LRE) 2022. The objective of this LRE (LRE22) is focused on recognizing 14 low-resourced African languages. Even though NIST has provided additional training and development data, we develop our systems with additional constraints of extreme low-resource. Our primary fixed-set submission ensures the usage of only the LRE 22 development data that contains the utterances of 14 target languages. We further restrict our system from using any pre-trained models for feature extraction or classifier fine-tuning. To address the issue of low-resource, our system relies on diverse audio augmentations followed by classifier fusions. Abiding by all the constraints, the proposed methods achieve an EER of 11.43% and cost metric of 0.41 in the LRE22 development set. For users with limited computational resources or limited storage/network capabilities, the proposed system will help achieve efficient LID performance.

语言识别低资源语音处理轻量部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。