arXiv:2409.17987cs.CVcs.HC2024-09被引 6

用大模型从脑影像还原视频语义,跨人差异小、效果好。

LLM4Brain: Training a Large Language Model for Brain Video Understanding

  • 用适配器微调脑影像编码器,将脑活动映射到与视频对齐的隐向量
  • 在多个语义指标上接近真实视频信息,跨被试泛化能力强
  • 适合神经科学、脑机接口研究者,可拓展至认知建模

从不同受试者的脑信号(如功能磁共振成像fMRI)中解码视觉-语义信息面临巨大挑战,包括信噪比低、数据有限和跨被试差异。近期大语言模型(LLM)在多模态处理中表现出色。本研究提出一种基于LLM的方法,从视频刺激诱发的fMRI信号中重建视觉-语义信息。具体地,通过在配备适配器的fMRI编码器上进行微调,将脑响应转化为与视频刺激对齐的潜在表示,再由LLM映射到文本模态。特别地,引入自监督领域适应方法,增强视觉-语义信息与脑反应之间的对齐。所提方法在多种定量语义度量下表现良好,生成结果与真实信息高度相似。

原文摘要 · Abstract (English)

Decoding visual-semantic information from brain signals, such as functional MRI (fMRI), across different subjects poses significant challenges, including low signal-to-noise ratio, limited data availability, and cross-subject variability. Recent advancements in large language models (LLMs) show remarkable effectiveness in processing multimodal information. In this study, we introduce an LLM-based approach for reconstructing visual-semantic information from fMRI signals elicited by video stimuli. Specifically, we employ fine-tuning techniques on an fMRI encoder equipped with adaptors to transform brain responses into latent representations aligned with the video stimuli. Subsequently, these representations are mapped to textual modality by LLM. In particular, we integrate self-supervised domain adaptation methods to enhance the alignment between visual-semantic information and brain responses. Our proposed method achieves good results using various quantitative semantic metrics, while yielding similarity with ground-truth information.

脑机接口大模型视觉语义fMRI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。