对比拼音与字符级输出,测试新型解码器在脑机文本生成中的表现。
Phoneme- vs. Character-Level Targets and Selective State-Space Models for Intracortical Brain-to-Text

- 用分层混合Mamba模型替代传统循环神经网络解码器
- 字符级目标下最优模型达13.39%的错误率(CER)
- 发现不同错误类型与大脑信号表征有关
当前先进的皮层脑-文本系统通常采用神经序列音素解码器配合外部语言模型。两个设计方向尚未充分探索:选择性状态空间模型(Mamba)是否优于循环解码器,以及输出目标(音素级或字符级)如何与解码器选择相互作用。在公开的Brain-to-Text '25基准上,我们基于统一可复现协议,构建了一个受控的2×2实验网格(GRU vs. 混合Mamba解码器;音素级 vs. 字符级目标),均使用CTC目标函数训练。循环基线依然最强:最佳音素级GRU达到12.62% PER和21.19% WER,最佳字符级GRU经语言模型重打分后达到13.39% CER和26.28% WER。混合Mamba模型表现良好但未超越。消融实验揭示了架构贡献,误差分析显示存在依赖表示的失败模式:发音相似的音素混淆,以及词汇和词边界错误。
原文摘要 · Abstract (English)
State-of-the-art intracortical brain-to-text systems pair a neural-sequence phone decoder with an external language model. Two design axes remain underexplored: whether selective state-space models (Mamba) improve on recurrent decoders, and how the output target (phonetic vs.\ character) interacts with that choice. On the public Brain-to-Text '25 benchmark, we study a controlled 2x2 grid (GRU vs.\ hybrid Mamba decoder; phonetic vs.\ character targets) trained with a CTC objective under one reproducible protocol. The recurrent baseline remains strongest: the best phonetic GRU reaches 12.62\% PER and 21.19\% WER, while the best textual GRU after LM rescoring reaches 13.39\% CER and 26.28\% WER. The Mamba hybrid is competitive but does not surpass it. Ablations isolate architectural contributions, and error analysis shows representation-dependent failures: articulatory-like phoneme confusions vs.\ lexical and word-boundary errors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。