arXiv:2409.14021cs.CVcs.AI2024-09被引 7

用脑电波和文字生成连贯图像,让想象直接变画面

BrainDreamer: Reasoning-Coherent and Controllable Image Generation from EEG Brain Signals via Language Guidance

  • 通过掩码对比学习对齐脑电信号、文本与图像特征
  • 在EEG数据噪声下仍生成高质量且逻辑一致的图像
  • 支持文字控制生成,适合脑机接口与创意设计场景

我们能否直接将脑海中的想象与语言描述共同可视化?人类感知的内在特性表明,思考时身体会结合语言描述构建生动的画面。直观地,生成模型也应具备这种能力。本文提出BrainDreamer,一种端到端的语言引导生成框架,可模仿人类推理,从脑电图(EEG)信号生成高质量图像。该方法在消除非侵入式EEG采集带来的噪声的同时,实现了脑电与图像模态之间更精确的映射,显著提升生成效果。具体而言,BrainDreamer包含两个关键阶段:1)模态对齐,提出基于掩码的三重对比学习策略,有效对齐EEG、文本与图像嵌入,学习统一表示;2)图像生成,通过设计可学习的EEG适配器,将EEG嵌入注入预训练Stable Diffusion模型,生成高质量且推理连贯的图像。此外,BrainDreamer可接受文本描述(如颜色、位置等)实现可控生成。大量实验表明,该方法在生成质量与定量指标上均显著优于现有方法。

原文摘要 · Abstract (English)

Can we directly visualize what we imagine in our brain together with what we describe? The inherent nature of human perception reveals that, when we think, our body can combine language description and build a vivid picture in our brain. Intuitively, generative models should also hold such versatility. In this paper, we introduce BrainDreamer, a novel end-to-end language-guided generative framework that can mimic human reasoning and generate high-quality images from electroencephalogram (EEG) brain signals. Our method is superior in its capacity to eliminate the noise introduced by non-invasive EEG data acquisition and meanwhile achieve a more precise mapping between the EEG and image modality, thus leading to significantly better-generated images. Specifically, BrainDreamer consists of two key learning stages: 1) modality alignment and 2) image generation. In the alignment stage, we propose a novel mask-based triple contrastive learning strategy to effectively align EEG, text, and image embeddings to learn a unified representation. In the generation stage, we inject the EEG embeddings into the pre-trained Stable Diffusion model by designing a learnable EEG adapter to generate high-quality reasoning-coherent images. Moreover, BrainDreamer can accept textual descriptions (e.g., color, position, etc.) to achieve controllable image generation. Extensive experiments show that our method significantly outperforms prior arts in terms of generating quality and quantitative performance.

脑机接口图像生成EEG可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。