用多模态模型让脑信号直接翻译成文字,效果比现有方法提升8.48%。
Decoding the Multimodal Mind: Generalizable Brain-to-Text Translation via Multimodal Alignment and Adaptive Routing
- 用多模态大模型对齐脑信号与文本、图像、音频的共同语义空间。
- 在多种fMRI数据上实现领先性能,基准测试提升8.48%。
- 可适配EEG/MEG,适合跨模态脑机接口实际应用。
从人脑解码语言仍是脑机接口(BCI)的重大挑战。现有方法通常依赖单一脑信号表征,忽视大脑固有的多模态处理机制。受大脑联想机制启发——看图可唤起相关声音和语言表征——我们提出一种统一框架,利用多模态大语言模型(MLLMs)将脑信号对齐至包含文本、图像和音频的共享语义空间。通过路由模块动态选择并融合不同模态的脑特征,以适应刺激特性。在多个包含文本、视觉和听觉刺激的fMRI数据集上实验表明,该方法达到当前最优性能,在最常用基准上提升8.48%。进一步扩展至EEG和MEG数据,验证了框架在不同时间和空间分辨率下的灵活性与鲁棒性。据我们所知,这是首个能跨多种脑信号与刺激类型稳健解码多模态脑活动的统一架构,为真实应用场景提供灵活解决方案。
原文摘要 · Abstract (English)
Decoding language from the human brain remains a grand challenge for Brain-Computer Interfaces (BCIs). Current approaches typically rely on unimodal brain representations, neglecting the brain's inherently multimodal processing. Inspired by the brain's associative mechanisms, where viewing an image can evoke related sounds and linguistic representations, we propose a unified framework that leverages Multimodal Large Language Models (MLLMs) to align brain signals with a shared semantic space encompassing text, images, and audio. A router module dynamically selects and fuses modality-specific brain features according to the characteristics of each stimulus. Experiments on various fMRI datasets with textual, visual, and auditory stimuli demonstrate state-of-the-art performance, achieving an 8.48% improvement on the most commonly used benchmark. We further extend our framework to EEG and MEG data, demonstrating flexibility and robustness across varying temporal and spatial resolutions. To our knowledge, this is the first unified BCI architecture capable of robustly decoding multimodal brain activity across diverse brain signals and stimulus types, offering a flexible solution for real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。