用少量受试者实现自然语音脑部解码,发现单人训练更优。
Optimizing fMRI Data Acquisition for Decoding Natural Speech with Limited Participants
- 用深度网络从fMRI预测大语言模型文本表征
- 8名受试者下,多人训练未提升解码准确率
- 复杂语法或丰富语义的故事更难解码
我们研究在受试者数量有限条件下,如何最优采集fMRI数据以解码感知到的自然语音。基于Lebel等人(2023)的8名受试者数据集,我们首先证明训练深度神经网络从fMRI活动预测LLM生成的文本表示是有效的。在该数据规模下,多人训练并未比单人训练提升解码准确率;跨被试使用相似或不同刺激训练对解码性能影响微乎其微。此外,我们的解码器更擅长建模句法特征而非语义特征,含有复杂句法或丰富语义内容的故事更难解码。结果表明,尽管每名被试的大量数据(深度表型)有益,但要利用多被试数据进行自然语音解码,仍需更深入的表型数据或更大样本量。
原文摘要 · Abstract (English)
We investigate optimal strategies for decoding perceived natural speech from fMRI data acquired from a limited number of participants. Leveraging Lebel et al. (2023)'s dataset of 8 participants, we first demonstrate the effectiveness of training deep neural networks to predict LLM-derived text representations from fMRI activity. Then, in this data regime, we observe that multi-subject training does not improve decoding accuracy compared to single-subject approach. Furthermore, training on similar or different stimuli across subjects has a negligible effect on decoding accuracy. Finally, we find that our decoders better model syntactic than semantic features, and that stories containing sentences with complex syntax or rich semantic content are more challenging to decode. While our results demonstrate the benefits of having extensive data per participant (deep phenotyping), they suggest that leveraging multi-subject for natural speech decoding likely requires deeper phenotyping or a substantially larger cohort.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。