首个阿拉伯语脑电语音识别数据集,助力残障人士沟通
ArEEG_Words: Dataset for Envisioned Speech Recognition using EEG for Arabic Words
- 用14通道设备采集22人想象16个阿拉伯词的脑电信号
- 共获352次记录,生成1.5万条250毫秒脑电片段
- 公开共享,填补阿拉伯语脑机接口研究数据空白
脑-机接口(BCI)旨在通过将神经信号转化为语言,帮助沟通障碍患者。其中,脑电图(EEG)用于测量大脑电活动。尽管在BCI EEG领域已取得显著进展,但仍存在重大局限:非英语语言(如阿拉伯语)公开可用的EEG数据集匮乏。为此,本文提出ArEEG_Words数据集,这是首个面向阿拉伯语的脑电图数据集。该数据集由22名参与者(平均年龄22岁,5名女性,17名男性)使用14通道Emotiv Epoc X设备采集,参与者在实验前8小时禁食咖啡、酒精、香烟等影响神经系统物质,并在安静环境中闭眼想象16个常用阿拉伯词汇(如上、下、左、右),每次想象持续10秒。共收集352次完整记录,每条记录被分割为多个250毫秒的信号片段,总计生成15,360条脑电信号。据我们所知,ArEEG_Words是首个针对阿拉伯语脑电语音识别的公开数据集,其开放共享有望推动该领域研究发展。
原文摘要 · Abstract (English)
Brain-Computer-Interface (BCI) aims to support communication-impaired patients by translating neural signals into speech. A notable research topic in BCI involves Electroencephalography (EEG) signals that measure the electrical activity in the brain. While significant advancements have been made in BCI EEG research, a major limitation still exists: the scarcity of publicly available EEG datasets for non-English languages, such as Arabic. To address this gap, we introduce in this paper ArEEG_Words dataset, a novel EEG dataset recorded from 22 participants with mean age of 22 years (5 female, 17 male) using a 14-channel Emotiv Epoc X device. The participants were asked to be free from any effects on their nervous system, such as coffee, alcohol, cigarettes, and so 8 hours before recording. They were asked to stay calm in a clam room during imagining one of the 16 Arabic Words for 10 seconds. The words include 16 commonly used words such as up, down, left, and right. A total of 352 EEG recordings were collected, then each recording was divided into multiple 250ms signals, resulting in a total of 15,360 EEG signals. To the best of our knowledge, ArEEG_Words data is the first of its kind in Arabic EEG domain. Moreover, it is publicly available for researchers as we hope that will fill the gap in Arabic EEG research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。