用游戏引擎生成教室语音数据,解决教育AI语音模型训练数据不足问题
SimClass: A Classroom Speech Dataset Generated via Game Engine Simulation For Automatic Speech Recognition Research
- 用游戏引擎模拟教室环境噪声,可扩展至其他场景
- 构建了含真实语音与噪声的合成数据集,接近真实教室表现
- 适合语音识别、语音增强研究者用于提升模型鲁棒性
教育领域中大规模教室语音数据的匮乏限制了AI语音模型的发展。现有公开教室数据集数量有限,且缺乏专用教室噪声语料库,导致无法使用标准数据增强技术。本文提出一种基于游戏引擎的可扩展教室噪声合成方法,构建了SimClass数据集,包含合成的教室噪声语料库和模拟的教室语音数据。语音数据通过将公开儿童语音语料与YouTube讲座视频配对生成,在清洁条件下模拟真实课堂互动。实验表明,该数据集在干净与嘈杂语音任务中均能逼近真实教室语音表现,为开发鲁棒的语音识别与增强模型提供了重要资源。
原文摘要 · Abstract (English)
The scarcity of large-scale classroom speech data has hindered the development of AI-driven speech models for education. Public classroom datasets remain limited, and the lack of a dedicated classroom noise corpus prevents the use of standard data augmentation techniques. In this paper, we introduce a scalable methodology for synthesizing classroom noise using game engines, a framework that extends to other domains. Using this methodology, we present SimClass, a dataset that includes both a synthesized classroom noise corpus and a simulated classroom speech dataset. The speech data is generated by pairing a public children's speech corpus with YouTube lecture videos to approximate real classroom interactions in clean conditions. Our experiments on clean and noisy speech demonstrate that SimClass closely approximates real classroom speech, making it a valuable resource for developing robust speech recognition and enhancement models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。