arXiv:2602.22597cs.SDeess.AS2026-02被引 1

研究三种说话方式的脑神经表征共性,发现线性模型可跨条件重建语音。

Relating the Neural Representations of Vocalized, Mimed, and Imagined Speech

  • 用线性模型分别训练三类说话模式的语音重建
  • 跨条件重建成功率高,说明神经表征有共享结构
  • 线性模型在细节还原上优于非线性网络

我们利用公开的立体定向脑电记录,研究了发声、模拟和想象说话的神经表征关系。以往研究多聚焦单一条件下的语音解码,本文则关注不同条件间的关联性:为每种条件训练线性频谱重建模型,并评估其跨条件泛化能力。结果表明,基于某一条件训练的线性解码器能有效应用于其他条件,暗示共享的语音神经表征存在。通过基于排序的分析验证了刺激级判别性在各条件间均保持稳定。最后,将线性模型与非线性神经网络的重建效果对比,两者均有跨条件迁移能力,但线性模型在刺激级判别性上表现更优。

原文摘要 · Abstract (English)

We investigated the relationship among neural representations of vocalized, mimed, and imagined speech recorded using publicly available stereotactic EEG recordings. Most prior studies have focused on decoding speech responses within each condition separately. Here, instead, we explore how responses across conditions relate by training linear spectrogram reconstruction models for each condition and evaluate their generalization across conditions. We demonstrate that linear decoders trained on one condition generally transfer successfully to others, implying shared speech representations. This commonality was assessed with stimulus-level discriminability by performing a rank-based analysis demonstrating preservation of stimulus-specific structure in both within- and across-conditions. Finally, we compared linear reconstructions to those from a nonlinear neural network. While both exhibited cross-condition transfer, linear models achieve superior stimulus-level discriminability.

神经表征语音生成脑机接口

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。