用图像生成脑电信号,为视觉假体提供完整闭环方案
Image-to-Brain Signal Generation for Visual Prosthesis with CLIP Guided Multimodal Diffusion Models
- 基于扩散变换器与跨注意力机制,将图像转为脑电信号
- 在两个数据集上生成的信号符合生物合理性,验证有效性
- 适合神经工程、脑机接口及视觉假体研究者参考
视觉假体有望恢复失明者的视力。尽管已有研究利用脑电/脑磁图(M/EEG)信号在解码阶段诱发视觉感知,但图像到脑信号的编码过程仍基本未被探索,阻碍了完整功能流程的建立。本文提出一种新型图像到脑信号生成框架,通过结合扩散变换器(DiT)与跨注意力机制,实现从图像生成M/EEG信号。我们采用基于去噪扩散隐式模型(DDIM)的扩散变换器架构,并引入跨注意力对齐脑信号嵌入与CLIP图像嵌入。进一步地,利用大语言模型(LLMs)生成图像描述,将对应的CLIP文本嵌入与图像嵌入拼接,形成统一嵌入用于对齐,以捕捉核心语义信息。同时,设计可学习的时空位置编码,融合脑区嵌入与时间嵌入,捕获脑信号的空间与时间特性。我们在两个多模态基准数据集(THINGS-EEG2 和 THINGS-MEG)上进行评估,结果表明生成的脑信号具有生物学合理性。
原文摘要 · Abstract (English)
Visual prostheses hold great promise for restoring vision in blind individuals. While researchers have successfully utilized M/EEG signals to evoke visual perceptions during the brain decoding stage of visual prostheses, the complementary process of converting images into M/EEG signals in the brain encoding stage remains largely unexplored, hindering the formation of a complete functional pipeline. In this work, we present a novel image-to-brain signal framework that generates M/EEG from images by leveraging the diffusion transformer architecture enhanced with cross-attention mechanisms. Specifically, we employ a diffusion transformer (DiT) architecture based on denoising diffusion implicit models (DDIM) to achieve brain signal generation. To realize the goal of image-to-brain signal conversion, we use cross-attention mechanisms to align brain signal embeddings with CLIP image embeddings. Moreover, we leverage large language models (LLMs) to generate image captions, and concatenate the resulting CLIP text embeddings with CLIP image embeddings to form unified embeddings for cross-attention alignment, enabling our model to capture core semantic information. Moreover, to capture core semantic information, we use large language models (LLMs) to generate descriptive and semantically accurate captions for images. Furthermore, we introduce a learnable spatio-temporal position encoding that combines brain region embeddings with temporal embeddings to capture both spatial and temporal characteristics of brain signals. We evaluate the framework on two multimodal benchmark datasets (THINGS-EEG2 and THINGS-MEG) and demonstrate that it generates biologically plausible brain signals.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。