arXiv:2603.20246cs.CLcs.AI2026-03

用上下文序列模型提升脑机接口语音解码精度与稳定性

Decoding the decoder: Contextual sequence-to-sequence modeling for intracortical speech decoding

  • 多任务Transformer联合预测音素、单词和声学特征
  • 音素错误率降至14.3%,单词错误率低至19.4%
  • 新校准模块显著抗日间波动,适合神经解码研究者

语音脑-机接口需要解码器将皮层电活动转化为语言输出,且在数据有限和日间变化下仍保持稳健。现有高性能系统多依赖逐帧音素解码结合下游语言模型,但上下文序列解码对子音素神经读取、鲁棒性与可解释性的贡献尚不明确。本文评估了一种基于Transformer的多任务序列到序列模型,用于从6v区皮层记录中解码尝试发声。该模型联合预测音素序列、词序列及辅助声学特征。为应对日间非平稳性,引入神经锤刀(NHS)校准模块,结合全局对齐与特征级调制。进一步分析了跨日泛化能力与编码器/解码器注意力模式。在Willett等人数据集上,音素错误率达14.3%的最新水平;直接解码词错误率为25.6%,候选生成与重打分后降至19.4%。NHS显著优于线性或无日特定变换;跨日实验显示,未见日期的性能随时间距离增加而下降。注意力可视化揭示编码器表示中存在重复的时间分块,音素与词解码器分别以不同方式利用这些片段。结果表明,上下文序列建模可提升皮层语音信号到音素读取的保真度,并提示注意力分析可生成关于神经语音证据如何分段累积的合理假设。

原文摘要 · Abstract (English)

Speech brain--computer interfaces require decoders that translate intracortical activity into linguistic output while remaining robust to limited data and day-to-day variability. While prior high-performing systems have largely relied on framewise phoneme decoding combined with downstream language models, it remains unclear what contextual sequence-to-sequence decoding contributes to sublexical neural readout, robustness, and interpretability. We evaluated a multitask Transformer-based sequence-to-sequence model for attempted speech decoding from area 6v intracortical recordings. The model jointly predicts phoneme sequences, word sequences, and auxiliary acoustic features. To address day-to-day nonstationarity, we introduced the Neural Hammer Scalpel (NHS) calibration module, which combines global alignment with feature-wise modulation. We further analyzed held-out-day generalization and attention patterns in the encoder and decoders. On the Willett et al. dataset, the proposed model achieved a state-of-the-art phoneme error rate of 14.3%. Word decoding reached 25.6% WER with direct decoding and 19.4% WER with candidate generation and rescoring. NHS substantially improved both phoneme and word decoding relative to linear or no day-specific transform, while held-out-day experiments showed increasing degradation on unseen days with temporal distance. Attention visualizations revealed recurring temporal chunking in encoder representations and distinct use of these segments by phoneme and word decoders. These results indicate that contextual sequence-to-sequence modeling can improve the fidelity of neural-to-phoneme readout from intracortical speech signals and suggest that attention-based analyses can generate useful hypotheses about how neural speech evidence is segmented and accumulated over time.

脑机接口语音解码Transformer神经解码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。