arXiv:2503.06437cs.CVcs.LG2025-03

提出新评估指标SEED,更准确衡量脑成像视觉解码的语义还原能力。

SEED: Towards More Accurate Semantic Evaluation for Visual Brain Decoding

  • 融合三种神经科学启发的语义相似度指标,综合评估解码效果。
  • 在人工评估数据上表现最优,现有模型仍丢失关键语义信息。
  • 开源人类评估数据集,助力脑解码评估方法研究。

我们提出SEED(语义评估用于视觉脑解码),一种新型视觉脑解码模型语义解码性能评估指标。该指标整合了三个互补的度量方式,分别借鉴神经科学发现,捕捉图像间的不同层面语义相似性。基于精心收集的人类标注数据,我们证明SEED与人类评估结果具有最高一致性,显著优于现有广泛使用的指标。利用SEED评估现有视觉脑解码模型后发现,即使在当前顶尖模型中,语义信息也常在解码过程中严重丢失,尽管其在原有指标上已接近完美得分。这一发现揭示了当前评估方法的局限性,并为未来模型改进提供方向。最后,为促进后续研究,我们公开了人类评估数据集,鼓励开发更先进的脑解码评估方法。代码与数据详见 https://github.com/Concarne2/SEED。

原文摘要 · Abstract (English)

We present SEED (Semantic Evaluation for Visual Brain Decoding), a novel metric for evaluating the semantic decoding performance of visual brain decoding models. It integrates three complementary metrics, each capturing a different aspect of semantic similarity between images inspired by neuroscientific findings. Using carefully crowd-sourced human evaluation data, we demonstrate that SEED achieves the highest alignment with human evaluation, outperforming other widely used metrics. Through the evaluation of existing visual brain decoding models with SEED, we further reveal that crucial information is often lost in translation, even in the state-of-the-art models that achieve near-perfect scores on existing metrics. This finding highlights the limitations of current evaluation practices and provides guidance for future improvements in decoding models. Finally, to facilitate further research, we open-source the human evaluation data, encouraging the development of more advanced evaluation methods for brain decoding. Our code and the human evaluation data are available at https://github.com/Concarne2/SEED.

脑解码语义评估人类标注开源数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。