首个融合音频、视觉与肌电的中文语音识别数据集,助力听障人群沟通。
AVE Speech: A Comprehensive Multi-Modal Dataset for Speech Recognition Integrating Audio, Visual, and Electromyographic Signals
- 采集100人10次朗读100句普通话,生成多模态数据
- 三模态融合使跨人识别和噪声环境下识别率显著提升
- 公开可用,适合语音识别与人机交互研究者使用
全球老龄化带来听力与发音障碍问题,影响沟通。为此,我们推出AVE Speech数据集,用于语音识别任务。该数据集包含100个普通话句子,涵盖音频信号、唇部视频和六通道肌电(EMG)数据,来自100名参与者。每位受试者重复朗读全部语料10次,每句平均约2秒,各模态数据总量超55小时。实验表明,多模态融合显著提升识别性能,尤其在跨主体及高噪声环境。据我们所知,这是首个公开的、面向大规模普通话语音识别的句子级三模态数据集。该数据集有望推动声学与非声学语音识别研究,促进跨模态学习与人机交互发展。
原文摘要 · Abstract (English)
The global aging population faces considerable challenges, particularly in communication, due to the prevalence of hearing and speech impairments. To address these, we introduce the AVE speech, a comprehensive multi-modal dataset for speech recognition tasks. The dataset includes a 100-sentence Mandarin corpus with audio signals, lip-region video recordings, and six-channel electromyography (EMG) data, collected from 100 participants. Each subject read the entire corpus ten times, with each sentence averaging approximately two seconds in duration, resulting in over 55 hours of multi-modal speech data per modality. Experiments demonstrate that combining these modalities significantly improves recognition performance, particularly in cross-subject and high-noise environments. To our knowledge, this is the first publicly available sentence-level dataset integrating these three modalities for large-scale Mandarin speech recognition. We expect this dataset to drive advancements in both acoustic and non-acoustic speech recognition research, enhancing cross-modal learning and human-machine interaction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。