用深度状态空间模型从脑影像解码语音可懂度,发现大脑有跨环境的抽象语言编码。
Condition-Invariant fMRI Decoding of Speech Intelligibility with Deep State Space Model
- 基于深度状态空间模型,捕捉fMRI高维时序特征以解码语音可懂度
- 在不同听觉条件下解码性能显著优于传统方法,跨条件转移有效
- 揭示听觉、额叶和顶叶区域共同参与抽象语言表征,适合神经科学与脑机接口研究者
阐明语音可懂度的神经基础对计算神经科学和数字语音处理至关重要。近期神经成像研究显示,可懂度不仅影响声学刺激,还主要作用于颞上回和下额回的皮层活动。然而,以往研究多集中于清晰语音,尚不清楚大脑是否在多样听觉环境中使用不变的神经编码。为此,我们提出一种基于深度状态空间模型的新架构,专为解析fMRI信号的高维时序结构而设计。这是首次在声学差异条件下解码语音可懂度的研究,结果表明该方法显著优于经典方法。区域分析显示听觉、额叶和顶叶区域均有贡献,且跨条件迁移能力表明存在条件不变的神经编码,从而深化了对大脑抽象语言表征的理解。
原文摘要 · Abstract (English)
Clarifying the neural basis of speech intelligibility is critical for computational neuroscience and digital speech processing. Recent neuroimaging studies have shown that intelligibility modulates cortical activity beyond simple acoustics, primarily in the superior temporal and inferior frontal gyri. However, previous studies have been largely confined to clean speech, leaving it unclear whether the brain employs condition-invariant neural codes across diverse listening environments. To address this gap, we propose a novel architecture built upon a deep state space model for decoding intelligibility from fMRI signals, specifically tailored to their high-dimensional temporal structure. We present the first attempt to decode intelligibility across acoustically distinct conditions, showing our method significantly outperforms classical approaches. Furthermore, region-wise analysis highlights contributions from auditory, frontal, and parietal regions, and cross-condition transfer indicates the presence of condition-invariant neural codes, thereby advancing understanding of abstract linguistic representations in the brain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。