arXiv:2410.20916cs.CL2024-10被引 9

统一处理多种脑电信号,实现更精准的脑波转文本。

NeuGPT: Unified multi-modal Neural GPT

  • 用统一模型处理EEG、fMRI等多模态脑信号
  • 脑波转文字性能提升,BLEU-1达12.92,ROUGE-1F达13.06
  • 可生成模拟脑信号,适用于神经接口研究

本文提出NeuGPT,一种突破性的多模态神经语言生成模型,旨在整合分散的神经记录研究。传统上,脑电(EEG)、脑磁(MEG)、皮层电图(ECoG)、深部脑电(SEEG)、功能磁共振(fMRI)和近红外(fNIRS)数据分别独立分析。我们认识到跨模态协同与神经信号在不同实验条件下的适应潜力,因此开发了可兼容多种模态的统一模型。受自然语言处理、计算机视觉和语音处理中预训练大模型成功的启发,NeuGPT能处理多样化神经记录并对接语音与文本数据。模型主要聚焦于脑到文本解码,将BLEU-1得分从6.94提升至12.92,ROUGE-1F从6.93提升至13.06。同时,该模型还能模拟脑信号,可作为新型神经接口使用。代码已开源。

原文摘要 · Abstract (English)

This paper introduces NeuGPT, a groundbreaking multi-modal language generation model designed to harmonize the fragmented landscape of neural recording research. Traditionally, studies in the field have been compartmentalized by signal type, with EEG, MEG, ECoG, SEEG, fMRI, and fNIRS data being analyzed in isolation. Recognizing the untapped potential for cross-pollination and the adaptability of neural signals across varying experimental conditions, we set out to develop a unified model capable of interfacing with multiple modalities. Drawing inspiration from the success of pre-trained large models in NLP, computer vision, and speech processing, NeuGPT is architected to process a diverse array of neural recordings and interact with speech and text data. Our model mainly focus on brain-to-text decoding, improving SOTA from 6.94 to 12.92 on BLEU-1 and 6.93 to 13.06 on ROUGE-1F. It can also simulate brain signals, thereby serving as a novel neural interface. Code is available at \href{https://github.com/NeuSpeech/NeuGPT}{NeuSpeech/NeuGPT (https://github.com/NeuSpeech/NeuGPT) .}

脑机接口多模态语言生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。