用反事实思维提升脑图重建的语义准确性,避免生成错误内容。
From "What-If" to "What-Is": Counterfactual Thinking-Inspired Semantic Alignment for Visual Brain Decoding

- 通过反事实样本对比,让模型区分正确与相似但错误的视觉内容。
- 在自然场景数据集上,语义判别能力提升32%,重建更贴近真实描述。
- 适合研究脑科学、神经影像和生成模型的交叉领域学者。
视觉脑解码从fMRI等神经测量中重构人所感知的视觉内容,是研究视觉信息在大脑中表征的计算方法。近年来多模态表示与扩散先验提升了重建的视觉真实性,但生成结果可能包含错误物体、属性或关系,因强生成先验会填补未充分指定的内容。传统评估仅关注最终图像,难以发现此类语义错误。本文提出ConceptAlign,一种受反事实思维启发的语义对齐框架:将解码的视觉令牌投影至冻结文本嵌入空间,使其与真实描述对齐,同时远离保留场景但存在细微差异的近似替代方案。这些替代方案由LLM离线生成,仅修改一个关键对象、属性或关系。采用基于边距的目标函数,在不需推理时调用LLM的前提下,学习细粒度语义边界。我们构建了涵盖基础可辨性、反事实描述判别与表征几何的三层次语义评估体系。在Natural Scenes Dataset上的实验表明,ConceptAlign在MindEye2基础上显著提升重建指标、反事实语义判别与表征对齐性能。负样本对照、独立LLM与人工生成替代、人类评估均验证了监督的有效性与鲁棒性,尤其在细粒度冲突、少样本解码与跨被试结构中表现良好。
原文摘要 · Abstract (English)
Visual brain decoding reconstructs visual content perceived by a person from neural measurements such as fMRI, providing a computational approach to studying how visual information is represented in the brain. Recent multimodal representations and diffusion priors have improved reconstruction realism. However, visually plausible reconstructions may contain incorrect objects, attributes, or relations because a strong generative prior can complete content not sufficiently specified by the decoded representation. Conventional reconstruction metrics mainly assess the final image and may therefore obscure such semantic errors. We propose ConceptAlign, a counterfactual semantic alignment framework for visual brain decoding. ConceptAlign pools decoded visual tokens and projects them into a frozen text-embedding space, aligning the representation with the ground-truth caption while separating it from scene-preserving near-miss alternatives. Generated offline by an LLM, these alternatives modify one critical object, attribute, or relation while retaining the scene. A margin-based objective learns fine-grained semantic boundaries between the observed stimulus and plausible but incorrect interpretations without requiring LLM calls during inference. We introduce a systematic three-level semantic evaluation framework covering foundational discriminability, counterfactual description discrimination, and representational geometry. Experiments on the Natural Scenes Dataset show that ConceptAlign improves reconstruction measures, counterfactual semantic discrimination, and representational alignment over the MindEye2 backbone. Matched negative-source ablations, independent LLM and human-written alternatives, and human evaluation support the effectiveness and robustness of the supervision, with favorable patterns in fine-grained conflicts, limited-data decoding, and cross-subject structure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。