构建可动态评估语音识别鲁棒性的模块化数据集
MoDiCoL: A Modular Diagnostic Continual Learning Dataset for Robust Speech Recognition

- 设计模块化数据集,可控分离语言、说话人和声学环境因素
- 提出真实场景启发的持续学习课程,模拟模型渐进适应过程
- 验证三种持续学习策略在动态条件下的鲁棒性演化
现代自动语音识别(ASR)系统在标准基准上表现优异,但在实际应用中因录音条件、口音、言语障碍和噪声等分布偏移导致性能下降。现有数据集通常孤立处理这些因素,忽视了它们在真实场景中的共现。本文认为模型鲁棒性是一种动态发展的能力,提出 MoDiCoL——一个用于语言内容、说话人特征和声学环境可控分析的模块化诊断持续学习数据集。同时,设计一种受真实场景启发的持续学习课程,模拟模型增量更新,研究鲁棒性如何获得、迁移与遗忘。评估三种持续学习策略,揭示其在动态条件下的鲁棒性演变机制。
原文摘要 · Abstract (English)
Modern Automatic Speech Recognition (ASR) systems have made remarkable progress on standard benchmarks, yet performance gaps have emerged under real-world distribution shifts, caused by recording conditions, accents, speech impairments, and noise. Existing datasets and benchmarks typically isolate these factors, which overlooks their co-occurrence in real-world applications. In this paper, we argue that model robustness can be treated as a dynamic capability that continually develops, and we introduce MoDiCoL, a Modular Diagnostic Continual Learning dataset designed for controlled analysis of linguistic content, speaker characteristics, and acoustic environments. Furthermore, we propose a real-world-inspired continual learning curriculum to simulate incremental updates and study how robustness is acquired, transferred, and forgotten. We evaluate three continual learning strategies and provide detailed insights into robustness under evolving conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。