揭示脑控语音模型跨说话模式的因果表示机制
Mechanistic Interpretability of Brain-to-Speech Models Across Speech Modes
- 通过激活修补与因果追踪,解析模型内部跨模态表征
- 发现跨模式传递依赖于层特定紧凑子空间,非分散神经元
- 成果适用于脑机接口与神经解码研究者
脑控语音解码模型在发声、默语和想象说话中均表现良好,但其在不同说话模式间捕捉与传递信息的根本机制尚不明确。本文采用机制可解释性方法,对神经语音解码器的内部表征进行因果分析。通过跨模式激活修补和三模态插值,检验语音表征是否离散或连续;结合粗到细的因果追踪与因果擦除,识别出足以支持跨模态传输的局部因果结构。在神经元层面进行激活修补,发现少数非分散的神经元子集影响跨模态传输。结果表明,不同说话模式位于共享的连续因果流形上,跨模态传输由层特定紧凑子空间介导,而非广泛分布的活动。本研究为脑控语音模型中说话模式信息的组织与使用提供了因果解释,揭示了跨模式的分层且方向依赖的表征结构。
原文摘要 · Abstract (English)
Brain-to-speech decoding models demonstrate robust performance in vocalized, mimed, and imagined speech; yet, the fundamental mechanisms via which these models capture and transmit information across different speech modalities are less explored. In this work, we use mechanistic interpretability to causally investigate the internal representations of a neural speech decoder. We perform cross-mode activation patching of internal activations across speech modes, and use tri-modal interpolation to examine whether speech representations vary discretely or continuously. We use coarse-to-fine causal tracing and causal scrubbing to find localized causal structure, allowing us to find internal subspaces that are sufficient for cross-mode transfer. In order to determine how finely distributed these effects are within layers, we perform neuron-level activation patching. We discover that small but not distributed subsets of neurons, rather than isolated units, affect the cross-mode transfer. Our results show that speech modes lie on a shared continuous causal manifold, and cross-mode transfer is mediated by compact, layer-specific subspaces rather than diffuse activity. Together, our findings give a causal explanation for how speech modality information is organized and used in brain-to-speech decoding models, revealing hierarchical and direction-dependent representational structure across speech modes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。