arXiv:2410.14971cs.AIcs.CL2024-10ACL被引 9

用音视频模型解码脑信号,提升文本生成准确率与鲁棒性

BrainECHO: Semantic Brain Signal Decoding through Vector-Quantized Spectrogram Reconstruction for Whisper-Enhanced Text Generation

  • 通过离散自编码将脑电谱图转为有限离散表示,降低噪声干扰
  • 在冻结潜空间中对齐脑信号与语音谱图,使BLEU-4提升3.65%
  • 结合Whisper模型约束解码,实现高精度且不依赖教师强制的生成

当前脑电/脑磁到文本的解码系统存在三大缺陷:(1) 依赖教师强制方法,影响推理鲁棒性;(2) 易受会话特异性噪声影响,跨被试泛化差;(3) 脑信号与语言表征间存在错位,因预训练语言模型主导。为此,我们提出BrainECHO(基于向量量化谱图重建的语义脑信号解码,用于耳语增强文本生成),一个分阶段框架,通过解耦表征学习在EEG和MEG数据集上达到领先性能。具体包括三阶段:(1) 离散自编码,将连续梅尔谱图转化为高质量离散表示;(2) 冻结对齐,将脑信号嵌入映射至冻结潜空间中的对应梅尔谱图嵌入,通过向量量化重建有效过滤会话特异性噪声,使BLEU-4得分提升3.65%;(3) 约束解码微调,利用预训练Whisper模型进行音频到文本转换,在信号适配与知识保留间取得平衡,实现74%-89%的解码BLEU分数,且无需过度依赖教师强制。BrainECHO在句子、会话和被试无关条件下均表现稳健,通过高斯噪声测试,展现出提升基于语言的脑机接口的潜力。

原文摘要 · Abstract (English)

Current EEG/MEG-to-text decoding systems suffer from three key limitations: (1) reliance on teacher-forcing methods, which compromises robustness during inference, (2) sensitivity to session-specific noise, hindering generalization across subjects, and (3) misalignment between brain signals and linguistic representations due to pre-trained language model over-dominance. To overcome these challenges, we propose BrainECHO (Brain signal decoding via vEctor-quantized speCtrogram reconstruction for WHisper-enhanced text generatiOn), a multi-stage framework that employs decoupled representation learning to achieve state-of-the-art performance on both EEG and MEG datasets. Specifically, BrainECHO consists of three stages: (1) Discrete autoencoding, which transforms continuous Mel spectrograms into a finite set of high-quality discrete representations for subsequent stages. (2) Frozen alignment, where brain signal embeddings are mapped to corresponding Mel spectrogram embeddings in a frozen latent space, effectively filtering session-specific noise through vector-quantized reconstruction, yielding a 3.65% improvement in BLEU-4 score. (3) Constrained decoding fine-tuning, which leverages the pre-trained Whisper model for audio-to-text translation, balancing signal adaptation with knowledge preservation, and achieving 74%-89% decoding BLEU scores without excessive reliance on teacher forcing. BrainECHO demonstrates robustness across sentence, session, and subject-independent conditions, passing Gaussian noise tests and showcasing its potential for enhancing language-based brain-computer interfaces.

脑机接口语音生成向量量化Whisper

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。