构建首个可支持开放词汇的可穿戴无声语音数据集,用于提升静默语音识别性能。
SoniSpeech: A Large-Scale Open-Vocabulary Tri-Modal Dataset for Wearable Silent Speech Interfaces

- 使用声学传感眼镜采集超声回波、语音和视频三模态数据。
- 包含34小时、1.8万条语句,覆盖5356个词且全音素覆盖。
- 提供首个开放词汇无声语音识别基准,支持研究与应用开发。
可穿戴无声语音接口(SSIs)受限于小规模封闭词汇。现有实现大词汇量的方法需侵入式硬件。本文提出SoniSpeech,首个大规模、开放词汇、三模态可穿戴无声语音数据集,基于声学传感眼镜采集。数据包含34小时、18,000条语句,三种同步模态:超声回波轮廓、有声语音与前视视频,涵盖发声与无声两种模式。数据源自SODA对话数据集,提供当代英语对话内容,含5,356个唯一词汇及完整音素覆盖。基于CTC的ResNet-34基线模型在开放词汇无声语音识别任务上达到26.3%词错误率(WER),为该任务提供首个基准。数据集可通过https://doi.org/10.7298/xjjr-9m85 获取。
原文摘要 · Abstract (English)
Wearable silent speech interfaces (SSIs) are limited to small, closed vocabularies. Approaches achieving larger vocabularies require obtrusive hardware such as facial electrodes. We present SoniSpeech, the first large-scale, open-vocabulary, trimodal dataset for wearable SSI using acoustic-sensing eyewear. It contains 34 hours across 18,000 utterances with three synchronized modalities: ultrasound echo profiles, voiced audio, and frontal video, in both voiced and silent modes. The corpus draws from the SODA dialogue dataset, providing contemporary conversational English with 5,356 unique words and full phoneme coverage. A CTC-based ResNet-34 baseline achieves 26.3% word error rate (WER) on open-vocabulary silent speech recognition, the first benchmark for this task. Dataset is available at https://doi.org/10.7298/xjjr-9m85
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。