脑信号微调让语音模型更贴近人脑的语音处理阶段
Brain-tuned Speech Models Better Reflect Speech Processing Stages in the Brain
- 用脑电数据微调语音模型,使其更符合人脑处理语音的层级结构
- 微调后模型晚期层与语义脑区的匹配度显著提升
- 适合研究人脑语言机制或开发类脑语音系统的人看
预训练自监督语音模型在语音任务中表现优异,但其表征层次不反映人类语音处理的层级:中间层蕴含丰富语义,晚期层语义较弱。近期研究发现,通过人类脑记录进行脑信号微调(brain-tuning)可增强模型的语义理解能力。本文进一步考察脑微调模型在多级语音处理阶段的拟合程度。结果表明,脑微调模型的晚期层在与语义语言脑区的对齐上显著优于预训练模型。逐层探测显示,早期层仍专注低频声学特征,而晚期层则成为完成复杂高层任务的最佳表征。这些发现说明,脑微调模型不仅性能更优,还展现出从声学到语义的清晰层级处理过程,使其成为研究人类语音处理更理想的模型工具。
原文摘要 · Abstract (English)
Pretrained self-supervised speech models excel in speech tasks but do not reflect the hierarchy of human speech processing, as they encode rich semantics in middle layers and poor semantics in late layers. Recent work showed that brain-tuning (fine-tuning models using human brain recordings) improves speech models' semantic understanding. Here, we examine how well brain-tuned models further reflect the brain's intermediate stages of speech processing. We find that late layers of brain-tuned models substantially improve over pretrained models in their alignment with semantic language regions. Further layer-wise probing reveals that early layers remain dedicated to low-level acoustic features, while late layers become the best at complex high-level tasks. These findings show that brain-tuned models not only perform better but also exhibit a well-defined hierarchical processing going from acoustic to semantic representations, making them better model organisms for human speech processing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。