arXiv:2606.01264q-bio.NCcs.HC2026-06

1020小时日本语语音产生的多模态数据集,含脑电、肌电和语音。

A 1000-hour EEG-EMG-audio dataset of Japanese speech production

论文配图:A 1000-hour EEG-EMG-audio dataset of Japanese speech production
图 1 · 摘自论文原文
  • 同步采集3名母语者说话时的脑电、面部肌电和语音信号
  • 覆盖62-128通道,总计1020小时,带语音事件标注
  • 适用于语音解码、跨设备适应及脑电表征学习研究

我们呈现了一个多模态数据集,包含三位健康日语母语者在开放式词汇出声说话过程中同步记录的1020小时头皮脑电图(EEG)、面部肌电图(EMG)和语音音频。数据通过三种脑电系统采集——超高密度系统(g.Pangolin)和两种帽式系统(g.SCARABEO 和 eegosports),通道数范围为62-128,历时数月多次会话。每次会话均提供时间同步的EEG、面部EMG和音频,并附有语音事件标注与转录。尽管以语音解码为主要目标,该数据集也支持多模态信号处理、伪影建模、纵向与跨设备适应以及脑电表征学习。技术验证包括跨参与者、设备和任务的功率谱密度与事件相关电位分析,结果显示出预期的1/f谱特性、任务相关的α波段抑制以及时间锁定的诱发电位。数据集以脑成像数据结构(BIDS)格式通过OpenNeuro发布,并采用CC0许可,支持语音相关及更广泛的脑电研究。

原文摘要 · Abstract (English)

We present a multimodal dataset of 1020 hours of simultaneously recorded scalp electroencephalography (EEG), facial electromyography (EMG), and speech audio from three healthy native Japanese speakers during open-vocabulary overt speech. Recordings were acquired with three EEG systems-an ultra-high-density system (g.Pangolin) and two cap-type systems (g.SCARABEO and eegosports), spanning 62-128 channels-across many sessions over several months. Each session provides time-synchronized EEG, facial EMG, and audio, together with speech-event annotations and transcriptions. Although collected with speech decoding as a primary motivation, the dataset also supports work on multimodal signal processing, artifact modeling, longitudinal and cross-device adaptation, and EEG representation learning. Technical validation included power spectral density and event-related potential analyses across participants, devices, and tasks, which showed the expected 1/f spectral profile, task-related alpha-band attenuation, and time-locked evoked responses. The dataset is released in Brain Imaging Data Structure (BIDS) format via OpenNeuro under a CC0 waiver to support both speech-related and broader EEG research.

脑电多模态语音生成数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。