首个大规模中文耳语语音数据集及识别基线,助力隐私通信与特殊场景交互。
AISHELL6-whisper: A Chinese Mandarin Audio-visual Whisper Speech Dataset with Speech Recognition Baselines
- 构建30小时耳语与正常语音配对的音视频数据集,含同步面部视频。
- 耳语识别CER达4.13%,正常语音1.11%,超越wTIMIT基准。
- 适合语音隐私、医疗沟通及噪声环境下的语音研究者使用。
耳语语音识别在保障敏感通信隐私、为声带受限患者提供沟通桥梁以及在噪声敏感环境中实现离散交互方面至关重要。然而,中文耳语语音识别的发展受限于缺乏大规模数据集。本文提出AISHELL6-Whisper,一个大规模开源的中文字幕音视频耳语语音数据集,包含每类30小时的耳语与对应正常语音,并配有同步正面面部视频。此外,我们基于Whisper-Flamingo框架提出一种音视频语音识别(AVSR)基线,采用并行训练策略对齐不同语音类型嵌入,并引入投影层适配耳语的频谱特性。该模型在自建数据集测试集上,耳语识别字符错误率(CER)为4.13%,正常语音为1.11%,并在wTIMIT基准上达到新最优性能。数据集与基线代码已开源:https://zutm.github.io/AISHELL6-Whisper。
原文摘要 · Abstract (English)
Whisper speech recognition is crucial not only for ensuring privacy in sensitive communications but also for providing a critical communication bridge for patients under vocal restraint and enabling discrete interaction in noise-sensitive environments. The development of Chinese mandarin audio-visual whisper speech recognition is hindered by the lack of large-scale datasets. We present AISHELL6-Whisper, a large-scale open-source audio-visual whisper speech dataset, featuring 30 hours each of whisper speech and parallel normal speech, with synchronized frontal facial videos. Moreover, we propose an audio-visual speech recognition (AVSR) baseline based on the Whisper-Flamingo framework, which integrates a parallel training strategy to align embeddings across speech types, and employs a projection layer to adapt to whisper speech's spectral properties. The model achieves a Character Error Rate (CER) of 4.13% for whisper speech and 1.11% for normal speech in the test set of our dataset, and establishes new state-of-the-art results on the wTIMIT benchmark. The dataset and the AVSR baseline codes are open-sourced at https://zutm.github.io/AISHELL6-Whisper.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。