用类脑分层结构提升脑电与视频的细粒度对齐效果
Achieving Fine-grained Cross-modal Understanding through Brain-inspired Hierarchical Representation Learning
- 模仿人脑视觉通路,分两阶段实现跨模态对齐
- 在fMRI-视频检索中超越现有方法,提升显著
- 适合研究脑认知机制或跨模态对齐的学者
由于大脑表征的内在复杂性以及神经数据与视觉输入之间的模态差异,理解大脑对视觉刺激的响应仍然具有挑战性。现有方法主要依赖神经解码生成任务或简单相关性,无法反映大脑视觉处理的层次性和时序特性。为此,我们提出NeuroAlign框架,受人类视觉系统层次结构启发,实现细粒度的fMRI-视频对齐。该框架采用双阶段机制:第一阶段通过神经-时间对比学习(NTCL)实现全局语义理解,显式建模模态间的双向时序动态;第二阶段通过增强型向量量化实现细粒度模式匹配。NTCL通过模态间双向预测建模时间动态,DynaSyncMM-EMA方法实现动态多模态融合与自适应加权。实验表明,NeuroAlign在跨模态检索任务中显著优于现有方法,为理解视觉认知机制建立了新范式。
原文摘要 · Abstract (English)
Understanding neural responses to visual stimuli remains challenging due to the inherent complexity of brain representations and the modality gap between neural data and visual inputs. Existing methods, mainly based on reducing neural decoding to generation tasks or simple correlations, fail to reflect the hierarchical and temporal processes of visual processing in the brain. To address these limitations, we present NeuroAlign, a novel framework for fine-grained fMRI-video alignment inspired by the hierarchical organization of the human visual system. Our framework implements a two-stage mechanism that mirrors biological visual pathways: global semantic understanding through Neural-Temporal Contrastive Learning (NTCL) and fine-grained pattern matching through enhanced vector quantization. NTCL explicitly models temporal dynamics through bidirectional prediction between modalities, while our DynaSyncMM-EMA approach enables dynamic multi-modal fusion with adaptive weighting. Experiments demonstrate that NeuroAlign significantly outperforms existing methods in cross-modal retrieval tasks, establishing a new paradigm for understanding visual cognitive mechanisms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。