首个车载多通道多人中文语音数据集,助力复杂驾驶场景下语音识别研究。
AISHELL-5: The First Open-Source In-Car Multi-Channel Multi-Speaker Speech Dataset for Automatic Speech Diarization and Recognition
- 构建车载多通道多说话人语音数据,含远场与近场信号。
- 覆盖100+小时真实驾驶场景,40小时环境噪声用于模拟。
- 提供可复现基线系统,支持语音分离与识别一体化测试。
本文介绍AISHELL-5,首个开源的车载多通道多说话人中文自动语音识别(ASR)数据集。该数据集包含两部分:(1) 在电动车中采集的超过100小时多通道语音数据,涵盖60多个真实驾驶场景;音频由每扇车门上的麦克风捕捉四路远场信号,同时通过高保真耳机麦克风获取每位说话人的近场信号;(2) 40小时真实环境噪声录音,用于模拟车载语音数据。此外,我们还提供一个开源、可复现的基线系统,包含基于语音源分离的前端模型,从远场信号中提取各说话人纯净语音,并结合语音识别模块实现个体内容转录。实验表明,主流ASR模型在AISHELL-5上面临显著挑战。我们认为,AISHELL-5将推动复杂驾驶场景下ASR系统的研究,建立首个公开可用的车载ASR基准。
原文摘要 · Abstract (English)
This paper delineates AISHELL-5, the first open-source in-car multi-channel multi-speaker Mandarin automatic speech recognition (ASR) dataset. AISHLL-5 includes two parts: (1) over 100 hours of multi-channel speech data recorded in an electric vehicle across more than 60 real driving scenarios. This audio data consists of four far-field speech signals captured by microphones located on each car door, as well as near-field signals obtained from high-fidelity headset microphones worn by each speaker. (2) a collection of 40 hours of real-world environmental noise recordings, which supports the in-car speech data simulation. Moreover, we also provide an open-access, reproducible baseline system based on this dataset. This system features a speech frontend model that employs speech source separation to extract each speaker's clean speech from the far-field signals, along with a speech recognition module that accurately transcribes the content of each individual speaker. Experimental results demonstrate the challenges faced by various mainstream ASR models when evaluated on the AISHELL-5. We firmly believe the AISHELL-5 dataset will significantly advance the research on ASR systems under complex driving scenarios by establishing the first publicly available in-car ASR benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。