用多模态大模型解析脑信号,实现视觉细节精准还原
Exploring The Visual Feature Space for Multimodal Neural Decoding
- 从预训练视觉模型中选取最优特征空间,提升脑信号映射精度
- 提出多粒度脑细节理解基准,可评估物体、属性与关系的还原能力
- 零样本方法支持细粒度重建,适合神经解码与可解释性研究
脑信号的复杂性推动了多模态AI在脑信号与视觉、文本数据对齐中的应用,以实现可解释的描述。然而,现有研究多局限于粗粒度解读,缺乏对物体描述、位置、属性及其关系的详细信息,导致视觉解码时重建结果模糊不准确。为此,我们分析了多模态大语言模型(MLLMs)中预训练视觉组件的不同视觉特征空间,并提出一种零样本多模态脑解码方法,可与这些模型交互,在多个粒度层级上进行解码。为评估模型在脑信号中解码精细细节的能力,我们构建了多粒度脑细节理解基准(MG-BrainDub),包含两个关键任务:详细描述与显著问题回答,其指标聚焦于物体、属性和关系等关键视觉元素。该方法显著提升了神经解码精度,支持更准确的神经解码应用。代码将开源至 https://github.com/weihaox/VINDEX。
原文摘要 · Abstract (English)
The intrication of brain signals drives research that leverages multimodal AI to align brain modalities with visual and textual data for explainable descriptions. However, most existing studies are limited to coarse interpretations, lacking essential details on object descriptions, locations, attributes, and their relationships. This leads to imprecise and ambiguous reconstructions when using such cues for visual decoding. To address this, we analyze different choices of vision feature spaces from pre-trained visual components within Multimodal Large Language Models (MLLMs) and introduce a zero-shot multimodal brain decoding method that interacts with these models to decode across multiple levels of granularities. % To assess a model's ability to decode fine details from brain signals, we propose the Multi-Granularity Brain Detail Understanding Benchmark (MG-BrainDub). This benchmark includes two key tasks: detailed descriptions and salient question-answering, with metrics highlighting key visual elements like objects, attributes, and relationships. Our approach enhances neural decoding precision and supports more accurate neuro-decoding applications. Code will be available at https://github.com/weihaox/VINDEX.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。