首次实现非侵入式脑电转文字,性能远超此前水平。
Unlocking Non-Invasive Brain-to-Text
- 用大语言模型重打分,将单字预测扩展为完整词汇转写。
- 引入预测填充解决生僻词问题,词汇量显著提升。
- 首次实现跨数据集规模化训练,准确率提升2.1-2.3倍。
尽管侵入式脑电转文字(B2T)已取得显著进展,但非侵入式方法在标准指标上仍无法超越随机猜测,阻碍了无手术脑机接口在瘫痪患者中恢复沟通能力的实现。本文首次提出显著超越关键基线的非侵入式B2T结果,使BLEU得分较以往提高1.4至2.6倍。这一突破源于三项贡献:(1) 将近期词分类模型与大语言模型(LLM)重打分结合,将单字预测器转化为闭词汇表B2T系统;(2) 提出预测填空方法处理未登录词(OOV),大幅扩展有效词汇量;(3) 首次实现非侵入式B2T模型跨数据集规模化,推动深度学习大规模应用,准确率提升2.1至2.3倍。这些成果揭示了数据质量与词汇规模的关键作用,扫清了实现实用非侵入式B2T系统的主要障碍。
原文摘要 · Abstract (English)
Despite major advances in surgical brain-to-text (B2T), i.e. transcribing speech from invasive brain recordings, non-invasive alternatives have yet to surpass even chance on standard metrics. This remains a barrier to building a non-invasive brain-computer interface (BCI) capable of restoring communication in paralysed individuals without surgery. Here, we present the first non-invasive B2T result that significantly exceeds these critical baselines, raising BLEU by $1.4\mathrm{-}2.6\times$ over prior work. This result is driven by three contributions: (1) we extend recent word-classification models with LLM-based rescoring, transforming single-word predictors into closed-vocabulary B2T systems; (2) we introduce a predictive in-filling approach to handle out-of-vocabulary (OOV) words, substantially expanding the effective vocabulary; and (3) we demonstrate, for the first time, how to scale non-invasive B2T models across datasets, unlocking deep learning at scale and improving accuracy by $2.1\mathrm{-}2.3\times$. Through these contributions, we offer new insights into the roles of data quality and vocabulary size. Together, our results remove a major obstacle to realising practical non-invasive B2T systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。