用视觉语言模型评估脑电到图像重建的语义一致性,超越传统像素指标。
Lost in Visual Translation: A VLM-Assisted Perceptual-Semantic Coherence Framework for EEG-to-Image Reconstruction

- 引入多视觉语言模型提问机制,量化图像感知与语义一致性。
- 新指标BCS在语义一致性上误差仅0.082,相关性达0.850。
- 适合关注脑机接口图像重建真实语义的研究者使用。
EEG-to-image评价应区分视觉保真度与可恢复语义。然而,现有重建结果模糊、失真且细节不足,导致SSIM、LPIPS和CLIP对语义可恢复输出过度惩罚或奖励看似合理但错误的结果。我们分析了来自ATM、ENIGMA、BrainVis和DreamDiffusion的6,855对真实图像与重建图像,通过语义探针、描述严苛度与盲区率以及可控退化实验发现,像素级指标与语义一致性几乎无相关性,而表征指标混淆了感知与语义误差。为此,我们提出一种脑机接口感知-语义协同框架,利用四个视觉语言模型(VLM)通过结构化问题评估图像对,生成容错感知对齐分数(T-PAS)和容错语义对齐分数(T-SAS)。其共识被凝练为BCI-Coherence Score(BCS),在数据集上实现T-PAS MAE为0.079(r=0.700)、T-SAS MAE为0.082(r=0.850)。人工验证显示联合一致性判断高度可靠,科恩κ系数为0.882±0.174,克里彭多夫α为0.882,支持以语义可恢复性替代通用视觉相似性。代码与资源见https://sukt03.github.io/BCS/。
原文摘要 · Abstract (English)
EEG-to-image evaluation should distinguish visual fidelity from recoverable meaning. Yet EEG-derived reconstructions are blurry, distorted, and low-detail, causing SSIM, LPIPS, and CLIP to penalize semantically recoverable outputs or reward plausible but incorrect ones. We analyze 6,855 ground-truth/reconstruction pairs from ATM, ENIGMA, BrainVis, and DreamDiffusion using semantic probes, caption harshness and blind-spot rates, and controlled degradations. Pixel metrics show near-zero correlation with semantic consistency, while representation metrics conflate perceptual and semantic errors. We therefore introduce a BCI-aware framework in which four VLMs assess image pairs through structured questions, producing Tolerant Perceptual Alignment Scores (T-PAS) and Tolerant Semantic Alignment Scores (T-SAS). Their consensus is distilled into the BCI-Coherence Score (BCS), a compact evaluator achieving a T-PAS MAE of 0.079 (r = 0.700) and a T-SAS MAE of 0.082 (r = 0.850) on our data. Human validation shows highly reliable joint coherence judgments, with Cohen's kappa = 0.882 +/- 0.174 and Krippendorff's alpha = 0.882, supporting perceptual-semantic recoverability over generic visual similarity. Code and resources are available at https://sukt03.github.io/BCS/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。