用脑电+眼神输入实现每分钟150字的高精度文字输出
iPhoneme: Brain-to-Text Communication for ALS Using ConformerXL Decoding
- 融合改进的ConformerXL模型与眼球+静默发音输入
- 92.14%音素准确率,词准确率达73.39%,比之前快3%
- 支持实时运行,适合渐冻症患者日常沟通使用
针对约17.3万至23.25万名全球肌萎缩侧索硬化症(ALS)患者言语障碍问题,现有高性能脑机接口仅在22至31名患者中验证。本文提出iPhoneme系统,通过联合建模与交互设计解决神经解码精度和输入界面难题。该系统采用改进的ConformerXL架构(192.9M参数),结合多尺度膨胀卷积、双向GRU与预归一化技术,提升神经信号稳定性;引入眼球辅助音素输入界面,避免眼动追踪中的“弥达斯触碰”问题。采用6-gram音素语言模型(基于310万条序列训练)与WFST束搜索(束宽128),在包含256通道颅内脑电的T15数据集上(45个会话,8071次试验),实现92.14%音素准确率(7.86%音素错误率)、73.39%词准确率(26.61%词错误率),较此前最优水平提升约3%。系统在CPU上延迟仅180毫秒,可实现实时高精度脑到文本通信。
原文摘要 · Abstract (English)
Brain-computer interfaces (BCIs) for speech restoration hold transformative potential for the approximately 173,000--232,500 individuals worldwide with ALS-related dysarthria. Despite recent progress, high-performance speech BCIs have been demonstrated in only 22--31 patients globally, largely due to limitations in neural decoding accuracy and practical input interfaces. We present iPhoneme, a brain-to-text communication system that jointly addresses these challenges through integrated modeling and interaction design. The system combines a deep learning phoneme decoder based on a modified Conformer architecture (ConformerXL, 192.9M parameters) with a gaze-assisted phoneme input interface that mitigates the Midas touch problem in eye-tracking systems. The acoustic model incorporates a temporal prenet with multi-scale dilated convolutions and bidirectional GRU for neural jitter correction, temporal subsampling for CTC stability, and Pre-RMSNorm stabilization across 12 encoder blocks, trained with AdamW and cosine scheduling. On the interaction side, iPhoneme introduces a chorded gaze-plus-silent-speech paradigm that replaces dwell-time selection, enabling more efficient input. We evaluate the system on the T15 dataset (45 sessions, 8,071 trials) of 256-channel intracranial EEG from speech motor cortex regions. A 6-gram phoneme language model trained on 3.1M sequences, combined with WFST beam search (beam=128), achieves 92.14% phoneme accuracy (7.86% PER) and 73.39% word accuracy (26.61% WER), approximately 3% above prior state-of-the-art. The system operates on CPU with 180 ms latency, demonstrating real-time, high-accuracy brain-to-text communication for ALS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。