arXiv:2507.07993cs.CVcs.AI2025-07AAAI被引 2

提出多粒度评估框架,精准区分脑视觉解码模型优劣。

Multigranular Evaluation for Brain Visual Decoding

  • 构建分层结构指标,从前景到组件逐级匹配解码图像
  • 结合大语言模型提取物体属性关系,实现语义层面精细对比
  • 适用于神经科学验证,帮助研究人员选型与改进模型

现有脑视觉解码评估方法主要依赖粗粒度指标,难以区分模型差异,缺乏神经科学依据,也无法捕捉细粒度视觉区别。为此,我们提出BASIC框架,统一评估解码图像与真实图像在结构保真度、推断一致性及上下文连贯性上的表现。在结构层面,设计基于分割的分层指标,包括前景、语义、实例和组件掩码,通过掩码结构间的粒度感知对应进行评估。在语义层面,利用多模态大语言模型提取包含物体、属性与关系的结构化场景表示,实现细节丰富、可扩展且上下文相关的对比。我们在多个刺激-脑成像数据集上对多种视觉解码方法进行了基准测试。该框架提供了更敏感、可解释且全面的评估基础。

原文摘要 · Abstract (English)

Existing evaluation protocols for brain visual decoding predominantly rely on coarse metrics that obscure inter-model differences, lack neuroscientific foundation, and fail to capture fine-grained visual distinctions. To address these limitations, we introduce BASIC, a unified, multigranular evaluation framework that jointly quantifies structural fidelity, inferential alignment, and contextual coherence between decoded and ground-truth images. For the structural level, we introduce a hierarchical suite of segmentation-based metrics, including foreground, semantic, instance, and component masks, anchored in granularity-aware correspondence across mask structures. For the semantic level, we extract structured scene representations encompassing objects, attributes, and relationships using multimodal large language models, enabling detailed, scalable, and context-rich comparisons with ground-truth stimuli. We benchmark a diverse set of visual decoding methods across multiple stimulus-neuroimaging datasets within this unified evaluation framework. Together, these criteria provide a more discriminative, interpretable, and comprehensive foundation for evaluating brain visual decoding methods.

脑科学视觉解码多粒度评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。