arXiv:2502.08866cs.CL2025-02被引 10

用脑电反应微调语音模型,提升语义表征能力

BrainWavLM: Fine-tuning Speech Representations with Brain Responses to Language

  • 用低秩适配(LoRA)端到端微调WavLM模型,实现非线性映射
  • 全皮层微调使编码性能平均提升,且稳定性更强
  • 模型跨被试泛化,无需标注即增强语义表示

语音编码模型通过听觉表征预测人脑对口语刺激的响应。现有高性能模型多采用线性映射人工神经网络隐藏状态与脑数据,但线性限制可能影响效果。本文使用低秩适配(LoRA)对基于WavLM的编码模型进行端到端微调,构建名为BrainWavLM的模型。实验表明,在全皮层上微调可显著提升平均编码性能并增强稳定性;尽管低层级区域如听觉皮层(AC)性能下降,但针对性微调这些区域可恢复其表现,同时保留其他皮层的增益。微调模型在不同被试间具有良好泛化能力,说明其学习到了稳健的脑似语音表征。进一步通过训练线性探测器发现,脑数据增强了语音模型的语义表征,且无需任何显式标注。结果表明,脑反馈微调能生成当前最佳的语音编码模型,非线性方法有望弥合人工与生物语义表征间的差距。

原文摘要 · Abstract (English)

Speech encoding models use auditory representations to predict how the human brain responds to spoken language stimuli. Most performant encoding models linearly map the hidden states of artificial neural networks to brain data, but this linear restriction may limit their effectiveness. In this work, we use low-rank adaptation (LoRA) to fine-tune a WavLM-based encoding model end-to-end on a brain encoding objective, producing a model we name BrainWavLM. We show that fine-tuning across all of cortex improves average encoding performance with greater stability than without LoRA. This improvement comes at the expense of low-level regions like auditory cortex (AC), but selectively fine-tuning on these areas improves performance in AC, while largely retaining gains made in the rest of cortex. Fine-tuned models generalized across subjects, indicating that they learned robust brain-like representations of the speech stimuli. Finally, by training linear probes, we showed that the brain data strengthened semantic representations in the speech model without any explicit annotations. Our results demonstrate that brain fine-tuning produces best-in-class speech encoding models, and that non-linear methods have the potential to bridge the gap between artificial and biological representations of semantics.

语音编码脑机接口语义表征LoRA

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。