用一周脑电数据预训练,显著提升语音解码效果。
From Minutes to Days: Scaling Intracranial Speech Decoding with Supervised Pretraining
- 用长达一周的临床记录预训练模型,数据量扩大百倍以上。
- 模型在真实场景下解码准确率大幅提升,效果随数据增长呈对数线性提升。
- 揭示了跨天脑信号结构漂移问题,适合长期脑机接口研究者参考。
以往语音从脑活动解码依赖于短时、高度控制实验中收集的有限神经记录。本文引入新框架,利用患者临床监测期间获得的长达一周的颅内脑电与音频同步数据,使训练数据规模扩大超过两个数量级。基于该预训练,对比学习模型在性能上显著超越仅使用传统实验数据训练的模型,且性能提升随数据量增加呈对数线性关系。对学习表征的分析显示,尽管脑活动反映语音特征,但其全局结构在不同日期间存在显著漂移,凸显了建模跨日变异性的必要性。本方法为在真实生活与受控任务场景下实现可扩展的脑活动解码与建模开辟了新路径。
原文摘要 · Abstract (English)
Decoding speech from brain activity has typically relied on limited neural recordings collected during short and highly controlled experiments. Here, we introduce a framework to leverage week-long intracranial and audio recordings from patients undergoing clinical monitoring, effectively increasing the training dataset size by over two orders of magnitude. With this pretraining, our contrastive learning model substantially outperforms models trained solely on classic experimental data, with gains that scale log-linearly with dataset size. Analysis of the learned representations reveals that, while brain activity represents speech features, its global structure largely drifts across days, highlighting the need for models that explicitly account for cross-day variability. Overall, our approach opens a scalable path toward decoding and modeling brain representations in both real-life and controlled task settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。