让视觉模型主动适应脑电数据,提升跨模态对齐效果
Shrinking the Teacher: An Adaptive Teaching Paradigm for Asymmetric EEG-Vision Alignment
- 设计自适应教学范式,让视觉模型动态缩小知识结构以匹配脑电信号能力
- 在零样本脑图检索任务中达到60.2%准确率,领先前代方法9.8个百分点
- 适合研究脑机接口、跨模态对齐的学者,尤其关注不对称模态融合场景
从脑电(EEG)信号解码视觉特征是神经科学的核心挑战,当前主流方法依赖跨模态对齐。我们指出视觉与脑信号之间的关系本质不对称,存在两大关键差距:保真度差距(因EEG固有的噪声与信号退化,相比视觉的高保真特征)和语义差距(因EEG概念表征浅薄,相比视觉丰富的语义深度)。以往方法常忽略这种不对称性,将两者视为平等对齐对象,导致泛化性能差。为此,我们提出自适应教学范式,使‘教师’模态(视觉)在任务引导下动态收缩并调整其知识结构,将其语义密集特征适配‘学生’模态(EEG)的表达能力。我们通过ShrinkAdapter实现该范式,该模块采用无残差设计与瓶颈结构,简单高效。大量实验验证了该范式的合理性与有效性。本方法在零样本脑图检索任务中取得60.2%的top-1准确率,较此前最优方法提升9.8个百分点。本工作为不对称对齐提供了新视角:教师需主动收缩与适应,以弥合视觉-脑信号鸿沟。
原文摘要 · Abstract (English)
Decoding visual features from EEG signals is a central challenge in neuroscience, with cross-modal alignment as the dominant approach. We argue that the relationship between visual and brain modalities is fundamentally asymmetric, characterized by two critical gaps: a Fidelity Gap (stemming from EEG's inherent noise and signal degradation, vs. vision's high-fidelity features) and a Semantic Gap (arising from EEG's shallow conceptual representation, vs. vision's rich semantic depth). Previous methods often overlook this asymmetry, forcing alignment between the two modalities as if they were equal partners and thereby leading to poor generalization. To address this, we propose the adaptive teaching paradigm. This paradigm empowers the ``teacher" modality (vision) to dynamically shrink and adjust its knowledge structure under task guidance, tailoring its semantically dense features to match the ``student" modality (EEG)'s capacity. We implement this paradigm with the ShrinkAdapter, a simple yet effective module featuring a residual-free design and a bottleneck structure. Through extensive experiments, we validate the underlying rationale and effectiveness of our paradigm. Our method achieves a top-1 accuracy of 60.2\% on the zero-shot brain-to-image retrieval task, surpassing previous state-of-the-art methods by a margin of 9.8\%. Our work introduces a new perspective for asymmetric alignment: the teacher must shrink and adapt to bridge the vision-brain gap.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。