100小时高精度脑电语音解码数据集,助力非侵入式脑机接口研究
LibriBrain100: One Hundred Hours of Broad and Deep MEG Data for Neural Speech Decoding at Scale

- 构建百万小时级脑电语音数据集,聚焦单人深度与多人广度并重
- 在单词分类任务上达顶尖性能,验证数据质量与规模价值
- 支持开源下载与竞赛评测,适合脑机接口与神经解码研究者
我们推出LibriBrain100,一个专为可复现、标准化评估设计的大规模脑磁图(MEG)语音解码数据集。该数据集比原版扩大一倍以上,包含超过100小时高质量MEG数据,记录受试者聆听自然连续语音时的脑活动。其中约80小时来自单一受试者,创下单人深度神经数据新纪录——比次优数据集多8倍,较其他数据集多约80倍。为验证深度设计的价值,我们在单词分类基准上测试,使用现有解码模型取得当前最优表现,证明数据质量与大规模个体内数据的有效性。由于每人收集80小时不现实,我们还额外采集了32名受试者各约40分钟的数据。在同一基准下,展示多主体广度数据的优势:对预训练模型进行监督微调可显著弥补个体数据量不足。我们提供标准训练/验证/测试划分,通过开源Python库实现数据的便捷下载、可选预处理和主流深度学习框架加载。此外,数据集与评估系统将配套开放机器学习竞赛及公开排行榜,推动标准化评估。最终目标是加速实用非侵入式脑机接口发展,帮助重度瘫痪患者恢复交流能力。
原文摘要 · Abstract (English)
We introduce LibriBrain100, a large-scale MEG dataset for speech decoding designed from the ground up for reproducible, standardised evaluation. LibriBrain100 more than doubles the size of the original LibriBrain release, resulting in over 100 hours of high-quality MEG acquired while subjects listened to naturalistic continuous speech. With $\sim$80 hours from a single subject, LibriBrain100 sets a new record for deep, within-subject neural data (8$\times$ more than the next comparable dataset and roughly 80$\times$ more than other datasets). To demonstrate the payoff of this depth-first design, we evaluate on a word-classification benchmark---an increasingly well-established stepping stone towards the open challenge of noninvasive brain-to-text decoding. Using an existing decoding model, we achieve state-of-the-art performance---validating both the quality of the recordings and the value of within-subject data at scale. Because collecting 80 hours of data per user is impractical for real-world applications, we also collected $\sim$40 minutes of additional data from each of 32 subjects. Using the same word-classification benchmark, we demonstrate the value of broad multi-subject data: supervised finetuning of a pre-trained model can substantially compensate for limited per-subject data. We provide standard train, validation, and test splits, all reproducible through an open-sourced Python library that supports easy downloading, optional preprocessing, and data loading for common deep learning frameworks. In addition, the dataset and evaluation infrastructure are being released alongside an open machine-learning competition with a public leaderboard for standardised benchmarking. Ultimately, our hope is that LibriBrain100 will accelerate progress towards practical non-invasive brain-computer interfaces, capable of restoring communication to people living with severe paralysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。