arXiv:2510.04670cs.AI2025-10中稿 · ICASSP 2026被引 2

通过动态路由提升多模态脑响应建模的个体适应性

Improving Multimodal Brain Encoding Model with Dynamic Subject-awareness Routing

  • 设计可插拔的专家混合解码器,按个体特征动态分配计算资源
  • 在多被试数据上实现跨个体泛化性能提升,优于主流基线
  • 解码过程可解释,专家使用模式与内容类型高度相关

自然情景下的功能磁共振成像编码需应对多模态输入、融合方式变化及显著的个体差异。本文提出AFIRE(多模态脑响应编码的无偏框架),一种标准化时序对齐后融合特征的接口;以及MIND,一个可即插即用的专家混合解码器,采用基于被试感知的动态门控机制。端到端训练下,AFIRE将解码器与上游融合分离,MIND结合依赖标记的Top-K稀疏路由与被试先验,实现个性化专家调用而不损失通用性。在多种多模态主干网络和被试上的实验表明,该方法持续优于强基线,增强跨被试泛化能力,并揭示与内容类型相关的可解释专家模式。该框架为新编码器与数据集提供简单接入点,支持自然情景神经影像研究的稳健、即插即用式性能提升。

原文摘要 · Abstract (English)

Naturalistic fMRI encoding must handle multimodal inputs, shifting fusion styles, and pronounced inter-subject variability. We introduce AFIRE (Agnostic Framework for Multimodal fMRI Response Encoding), an agnostic interface that standardizes time-aligned post-fusion tokens from varied encoders, and MIND, a plug-and-play Mixture-of-Experts decoder with a subject-aware dynamic gating. Trained end-to-end for whole-brain prediction, AFIRE decouples the decoder from upstream fusion, while MIND combines token-dependent Top-K sparse routing with a subject prior to personalize expert usage without sacrificing generality. Experiments across multiple multimodal backbones and subjects show consistent improvements over strong baselines, enhanced cross-subject generalization, and interpretable expert patterns that correlate with content type. The framework offers a simple attachment point for new encoders and datasets, enabling robust, plug-and-improve performance for naturalistic neuroimaging studies.

脑编码多模态专家混合个体差异

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。