arXiv:2410.07507cs.CL2024-10NAACL被引 39

用脑电波直接生成文本,让思维变文字。

Thought2Text: Text Generation from EEG Signal using Large Language Models (LLMs)

  • 用大模型结合脑电数据,从脑信号生成文字
  • 在6名受试者上验证,能准确描述图像内容
  • 适合神经科学与脑机接口研究者使用

解码并以可理解形式表达脑活动是人工智能的前沿挑战。本文提出Thought2Text,通过指令微调的大语言模型(LLMs)结合脑电数据实现此目标。方法分三步:(1) 训练脑电编码器提取视觉特征;(2) 在图像与文本数据上微调大模型,实现多模态描述生成;(3) 在脑电嵌入上进一步微调,推理时直接由脑电生成文本。在包含六名受试者、图像刺激和对应文本标签的公开脑电数据集上进行实验,验证了多个大模型(LLaMA-v3、Mistral-v0.3、Qwen2.5)的有效性,采用传统语言生成评估指标及流畅性、充分性度量。该方法标志着低成本、便携式“脑电转文字”技术的重要进展,具有神经科学与自然语言处理双重应用潜力。

原文摘要 · Abstract (English)

Decoding and expressing brain activity in a comprehensible form is a challenging frontier in AI. This paper presents Thought2Text, which uses instruction-tuned Large Language Models (LLMs) fine-tuned with EEG data to achieve this goal. The approach involves three stages: (1) training an EEG encoder for visual feature extraction, (2) fine-tuning LLMs on image and text data, enabling multimodal description generation, and (3) further fine-tuning on EEG embeddings to generate text directly from EEG during inference. Experiments on a public EEG dataset collected for six subjects with image stimuli and text captions demonstrate the efficacy of multimodal LLMs (LLaMA-v3, Mistral-v0.3, Qwen2.5), validated using traditional language generation evaluation metrics, as well as fluency and adequacy measures. This approach marks a significant advancement towards portable, low-cost "thoughts-to-text" technology with potential applications in both neuroscience and natural language processing.

脑机接口文本生成EEG

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。