提出真实世界多语言多领域语音持续学习评估框架
NIRANTAR: Continual Learning with New Languages and Domains on Real-world Speech Data
- 基于印度22语言208地区真实数据构建增量学习场景
- 涵盖语言、领域及二者联合增量学习,共3250小时标注语音
- 揭示现有方法无一在所有场景稳定表现,凸显鲁棒性需求
我们提出Nirantar,一个用于多语言多领域自动语音识别持续学习(CL)的综合性评估框架。该框架基于印度22种语言、208个地区的自然增量语音数据,真实反映现实世界中持续学习的挑战。它支持语言增量(LIL)、领域增量(DIL)以及全新的语言-领域联合增量学习(LIDIL)三种场景。与以往依赖模拟数据的研究不同,Nirantar呈现动态、非均匀的语言和领域分布变化,是理想的研究测试平台。本工作包含3250小时人工转录语音,其中1720小时为新引入数据,可系统评估各类持续学习方法。实验表明,当前方法在不同场景下表现不一致,凸显开发更鲁棒策略的必要性。
原文摘要 · Abstract (English)
We introduce Nirantar, a comprehensive framework for evaluating continual learning (CL) in multilingual and multi-domain ASR. Designed to reflect real-world CL challenges, Nirantar leverages data collected incrementally across 22 languages and 208 districts in India through natural episodes. This enables evaluation across Language-Incremental (LIL), Domain-Incremental (DIL), and the novel Language-Incremental Domain-Incremental Learning (LIDIL) scenarios. Unlike prior work that relies on simulated episodes, Nirantar presents dynamic, non-uniform language and domain shifts, making it an ideal testbed for CL research. With 3250 hours of human-transcribed speech, including 1720 hours newly introduced in this work, our framework enables systematic benchmarking of CL methods. We evaluate existing approaches and demonstrate that no single method performs consistently well, underscoring the need for more robust CL strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。