arXiv:2410.15161cs.HCcs.CL2024-10被引 1

用大模型优化脑电拼写器,提升渐冻症患者沟通速度。

Evaluation Of P300 Speller Performance Using Large Language Models Along With Cross-Subject Training

  • 结合大语言模型动态优化字符提示和词预测
  • 对罕见词和生僻词的打字速度提升约40%
  • 跨被试训练下仍有效,适合临床辅助沟通

肌萎缩侧索硬化症(ALS)会严重限制患者数年内沟通能力,导致生活质量显著下降。基于P300的脑机接口(BCI)通过用户对界面字符网格中高亮字符的脑电信号响应,提供替代性交流方式。本研究针对多被试分类器训练局限及交互效率问题,引入GPT2、BERT、BART等大语言模型,结合Dijkstra算法优化刺激呈现与词补全,提升通信效率。采用多层平滑策略处理未登录词(OOV)。基于随机抽取的受试者脑电数据进行大规模仿真,结果显示:在包含稀有词和未登录词的文本输入中,字符级优化使打字速度提升约10%;使用GPT2进行多词预测时,速度提升达约40%。部分模型性能已接近本研究建立的理论上限(相差不足10%)。无论在同被试还是跨被试场景下,训练方法均显示显著提速效果。

原文摘要 · Abstract (English)

Amyotrophic lateral sclerosis (ALS), a progressive neuromuscular degenerative disease, severely restricts patient communication capacity within a few years of onset, resulting in a significant deterioration of quality of life. The P300 speller brain computer interface (BCI) offers an alternative communication medium by leveraging a subject's EEG response to characters traditionally highlighted on a character grid on a graphical user interface (GUI). A recurring theme in P300-based research is enhancing performance to enable faster subject interaction. This study builds on that theme by addressing key limitations, particularly in the training of multi-subject classifiers, and by integrating advanced language models to optimize stimuli presentation and word prediction, thereby improving communication efficiency. Furthermore, various advanced large language models such as Generative Pre-Trained Transformer (GPT2), BERT, and BART, alongside Dijkstra's algorithm, are utilized to optimize stimuli and provide word completion choices based on the spelling history. In addition, a multi-layered smoothing approach is applied to allow for out-of-vocabulary (OOV) words. By conducting extensive simulations based on randomly sampled EEG data from subjects, we show substantial speed improvements in typing passages that include rare and out-of-vocabulary (OOV) words, with the extent of improvement varying depending on the language model utilized. The gains through such character-level interface optimizations are approximately 10%, and GPT2 for multi-word prediction provides gains of around 40%. In particular, some large language models achieve performance levels within 10% of the theoretical performance limits established in this study. In addition, both within and across subjects, training techniques are explored, and speed improvements are shown to hold in both cases.

脑机接口大模型语音生成医疗辅助

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。