通过软脑区融合提升跨被试脑图像生成文本的准确性和可解释性。
Unified Multimodal Brain Decoding via Cross-Subject Soft-ROI Fusion
- 用软脑区作为共享空间,实现跨被试功能拓扑差异的统一建模。
- 在NSD数据集上,BLEU-4和CIDEr指标优于现有方法,显著提升跨被试性能。
- 引入可解释的提示优化流程,支持小样本闭环迭代与可审计设计。
多模态脑解码旨在从fMRI等脑活动信号中重建与视觉刺激一致的语义信息,并生成可读的自然语言描述。然而,该任务在跨被试泛化和可解释性方面仍面临挑战。本文提出BrainROI模型,在NSD数据集的脑图像生成文本评测中取得领先水平。在跨被试设置下,相比近期最先进方法和代表性基线,BLEU-4和CIDEr等指标均有明显提升。首先,为解决被试间功能脑拓扑异质性,设计新型fMRI编码器,采用多图谱软功能分区(soft-ROI)作为共享空间,并将MINDLLM中的离散脑区拼接扩展为体素级门控融合机制(Voxel-gate),结合全局标签对齐确保一致的脑区映射,增强跨被试迁移能力。其次,为克服人工与黑盒提示方法在稳定性和透明性上的局限,引入可解释的提示优化流程:在小样本闭循环中,利用本地部署的Qwen模型迭代生成并筛选人类可读提示,提升提示设计稳定性并保留可审计优化轨迹。最后,推理阶段施加参数化解码约束,进一步提升生成描述的稳定性和质量。
原文摘要 · Abstract (English)
Multimodal brain decoding aims to reconstruct semantic information that is consistent with visual stimuli from brain activity signals such as fMRI, and then generate readable natural language descriptions. However, multimodal brain decoding still faces key challenges in cross-subject generalization and interpretability. We propose a BrainROI model and achieve leading-level results in brain-captioning evaluation on the NSD dataset. Under the cross-subject setting, compared with recent state-of-the-art methods and representative baselines, metrics such as BLEU-4 and CIDEr show clear improvements. Firstly, to address the heterogeneity of functional brain topology across subjects, we design a new fMRI encoder. We use multi-atlas soft functional parcellations (soft-ROI) as a shared space. We extend the discrete ROI Concatenation strategy in MINDLLM to a voxel-wise gated fusion mechanism (Voxel-gate). We also ensure consistent ROI mapping through global label alignment, which enhances cross-subject transferability. Secondly, to overcome the limitations of manual and black-box prompting methods in stability and transparency, we introduce an interpretable prompt optimization process. In a small-sample closed loop, we use a locally deployed Qwen model to iteratively generate and select human-readable prompts. This process improves the stability of prompt design and preserves an auditable optimization trajectory. Finally, we impose parameterized decoding constraints during inference to further improve the stability and quality of the generated descriptions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。