arXiv:2604.02362cs.CLcs.AI2026-04被引 1

用脑电数据解码发音,提出双路径模型提升识别精度。

CIPHER: Conformer-based Inference of Phonemes from High-density EEG

论文配图:CIPHER: Conformer-based Inference of Phonemes from High-density EEG
图 1 · 摘自论文原文
  • 结合脑电反应与宽带频谱特征的双路径模型
  • 11类辅音-元音-辅音发音识别错误率达67%以上
  • 强调实验设计对干扰因素的控制,适合作为基准研究

从头皮脑电(EEG)中解码语音信息仍面临信噪比低和空间模糊的挑战。本文提出CIPHER(基于Conformer的高密度脑电发音推断模型),采用双路径架构,分别利用事件相关电位(ERP)特征和宽带动态谱分析(DDA)系数。在OpenNeuro ds006104数据集(24名参与者,含经颅磁刺激同步实验)上,二分类发音任务接近完美性能,但易受声学起始分离性和磁刺激靶点阻断等混杂因素影响。在主任务11类辅音-元音-辅音(CVC)发音识别中,全研究2留一被试交叉验证(LOSO)下表现显著降低:真实词错误率(WER)为ERP 0.671 ± 0.080,DDA 0.688 ± 0.096,表明精细发音区分能力有限。因此,本工作定位为基准与特征对比研究,而非端到端脑电转文本系统,并将神经表征结论严格限定于混杂因素可控的证据。

原文摘要 · Abstract (English)

Decoding speech information from scalp EEG remains difficult due to low SNR and spatial blurring. We present CIPHER (Conformer-based Inference of Phonemes from High-density EEG Representations), a dual-pathway model using (i) ERP features and (ii) broadband DDA coefficients. On OpenNeuro ds006104 (24 participants, two studies with concurrent TMS), binary articulatory tasks reach near-ceiling performance but are highly confound-vulnerable (acoustic onset separability and TMS-target blocking). On the primary 11-class CVC phoneme task under full Study 2 LOSO (16 held-out subjects), performance is substantially lower (real-word WER: ERP 0.671 +/- 0.080, DDA 0.688 +/- 0.096, indicating limited fine-grained discriminability. We therefore position this work as a benchmark and feature-comparison study rather than an EEG-to-text system, and we constrain neural-representation claims to confound-controlled evidence.

脑机接口语音解码EEG分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。