arXiv:2410.00712q-bio.NCcs.LG2024-10被引 15

用脑电波生成图像,突破分类限制,首次实现零样本图像生成。

NECOMIMI: Neural-Cognitive Multimodal EEG-informed Image Generation with Diffusion Models

  • 基于扩散模型构建脑电-图像生成框架,引入新型NERV编码器。
  • 在200类零样本任务中达当前最佳,生成图像语义质量显著提升。
  • 提出新评估指标CAT Score,适合研究脑机接口与生成模型交叉领域。

NECOMIMI(神经认知多模态脑电信息图像生成)提出一种基于扩散模型的全新框架,直接从脑电(EEG)信号生成图像。不同于以往仅聚焦于脑电-图像分类的对比学习方法,该工作首次将任务拓展至图像生成。所提出的NERV EEG编码器在多个零样本分类任务中表现优异,包括二分类、四分类及200类任务,并在新提出的类别评估表(CAT Score)中取得领先结果,该指标用于评估脑电生成图像的语义质量。关键发现是模型更倾向于生成抽象或泛化图像(如风景),而非具体物体,揭示了噪声大、分辨率低的脑电信号难以生成细节丰富视觉内容的根本挑战。研究还引入了针对脑电到图像生成的评估标准,在ThingsEEG数据集上建立了基准。本工作展示了脑电到图像生成的潜力,同时揭示了其背后仍存在的复杂性与技术瓶颈。

原文摘要 · Abstract (English)

NECOMIMI (NEural-COgnitive MultImodal EEG-Informed Image Generation with Diffusion Models) introduces a novel framework for generating images directly from EEG signals using advanced diffusion models. Unlike previous works that focused solely on EEG-image classification through contrastive learning, NECOMIMI extends this task to image generation. The proposed NERV EEG encoder demonstrates state-of-the-art (SoTA) performance across multiple zero-shot classification tasks, including 2-way, 4-way, and 200-way, and achieves top results in our newly proposed Category-based Assessment Table (CAT) Score, which evaluates the quality of EEG-generated images based on semantic concepts. A key discovery of this work is that the model tends to generate abstract or generalized images, such as landscapes, rather than specific objects, highlighting the inherent challenges of translating noisy and low-resolution EEG data into detailed visual outputs. Additionally, we introduce the CAT Score as a new metric tailored for EEG-to-image evaluation and establish a benchmark on the ThingsEEG dataset. This study underscores the potential of EEG-to-image generation while revealing the complexities and challenges that remain in bridging neural activity with visual representation.

脑机接口扩散模型图像生成多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。