用双曲空间插值脑信号与视觉特征,提升脑图对齐精度。
HyFI: Hyperbolic Feature Interpolation for Brain-Vision Alignment
- 在双曲空间中沿测地线插值语义与感知特征
- 在THINGS-EEG上实现17.3%的检索准确率提升
- 适合研究脑机接口与视觉认知的学者
近年来人工智能进展推动了从脑信号理解与解码人类视觉系统的研究。以往方法通常独立地将神经活动与预训练视觉模型提取的语义和感知特征对齐,但未能解决两大挑战:(1)脑信号与图像间因表征信息层级差异导致的模态鸿沟;(2)神经活动中语义与感知特征高度纠缠。为此,本文利用双曲空间——其天然适合处理信息量差异,并具备测地线向原点弯曲的几何特性(原点代表低表达能力区域)。基于此,提出新型框架HyFI,通过双曲测地线插值语义与感知视觉特征,实现感知与语义信息的融合与压缩,有效反映脑信号的表达限制及特征纠缠性。实验表明,该方法在零样本脑到图像检索任务中达到当前最优性能,在THINGS-EEG上Top-1准确率提升达+17.3%,在THINGS-MEG上提升+9.1%。
原文摘要 · Abstract (English)
Recent progress in artificial intelligence has encouraged numerous attempts to understand and decode human visual system from brain signals. These prior works typically align neural activity independently with semantic and perceptual features extracted from images using pre-trained vision models. However, they fail to account for two key challenges: (1) the modality gap arising from the natural difference in the information level of representation between brain signals and images, and (2) the fact that semantic and perceptual features are highly entangled within neural activity. To address these issues, we utilize hyperbolic space, which is well-suited for considering differences in the amount of information and has the geometric property that geodesics between two points naturally bend toward the origin, where the representational capacity is lower. Leveraging these properties, we propose a novel framework, Hyperbolic Feature Interpolation (HyFI), which interpolates between semantic and perceptual visual features along hyperbolic geodesics. This enables both the fusion and compression of perceptual and semantic information, effectively reflecting the limited expressiveness of brain signals and the entangled nature of these features. As a result, it facilitates better alignment between brain and visual features. We demonstrate that HyFI achieves state-of-the-art performance in zero-shot brain-to-image retrieval, outperforming prior methods with Top-1 accuracy improvements of up to +17.3% on THINGS-EEG and +9.1% on THINGS-MEG.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。