arXiv:2607.12071cs.CL2026-07

通过交互式多特征融合,提升非侵入脑信号的语义重建效果

Beyond Parallel Tracking: Interactive Multi-Feature Fusion Drives Semantic Reconstruction from Non-invasive Brain Recordings

论文配图:Beyond Parallel Tracking: Interactive Multi-Feature Fusion Drives Semantic Reconstruction from Non-invasive Brain Recordings
图 1 · 摘自论文原文
  • 融合静态词向量与动态上下文表示,用门控机制实现协同处理
  • 交叉注意力融合方法性能最优,显著优于串联和单一特征方法
  • 适合脑机接口、神经语言解码研究者参考

从非侵入性神经记录中实现连续语义重建仍受限于语义特征空间与神经编码模式之间的表征不匹配,严重阻碍了高噪声神经信号与目标语义特征间的跨模态对齐。以往语义解码器主要依赖静态词汇表示或动态上下文表示中的单一维度,导致信息损失。本文提出一种多特征融合框架,系统评估线性拼接与非线性多头交叉注意力两种集成方式。通过交互式门控机制,将静态词向量(W2V)与动态上下文表示(GPT)融合,促进语言理解过程中的协同处理。在广泛的语义重建与文本生成实验中,性能表现呈现显著层次:交叉注意力 > 拼接 > GPT > W2V。关键发现为:非线性交叉注意力融合方法达到当前最优水平,表明神经语言解码需模拟上下文信息与核心词汇属性间的协同调制,而非依赖孤立特征;同时提供了一种可行的非侵入式脑到文本解码路径。

原文摘要 · Abstract (English)

Continuous semantic reconstruction from non-invasive neural recordings remains limited by the representational mismatch between semantic feature spaces and neural coding patterns, which severely impedes cross-modal alignment between high-noise neural signals and target semantic features. Prior semantic decoders have predominantly relied on static lexical representations or dynamic contextualized representations in isolation. This single-dimension approach inevitably leads to severe information loss, as it fails to account for the human brain's capacity to integrate stable word attributes and dynamic contexts simultaneously. To bridge this gap, this study introduces a multi-feature fusion framework for non-invasive semantic reconstruction, systematically benchmarking two integration approaches: linear Naive Concatenation and non-linear Multi-Head Cross-Attention. Within this framework, our approach complements static lexical representations (W2V) with dynamic contextual representations (GPT) via an interactive gating mechanism to facilitate cooperative processing during language comprehension. Evaluated through extensive semantic reconstruction and text generation experiments, our framework reveals a robust performance hierarchy: Cross-Att > Concat > GPT > W2V. Crucially, the non-linear cross-attention fusion method achieves state-of-the-art performance, demonstrating that neural language decoding benefits from simulating the collaborative modulation between contextual information and core lexical attributes rather than depending on isolated individual features, while also offering a viable non-invasive brain-to-text decoding method.

脑机接口语义重建多特征融合交叉注意力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。