让大脑信号与语言语义对齐,提升脑-语言解码的可解释性。
Margin-Regularized Structured Semantic Alignment for Brain-Language Correspondence

- 通过边缘正则化对齐大脑与文本嵌入,显式建模神经表征与语义关系。
- 在全词汇和子集评估中均达到顶尖检索性能,验证方法有效性。
- 适合关注脑-语言机制、可解释性解码的研究者使用。
随着大语言模型的快速发展,脑-语言解码取得了显著进展。然而,解码内容是否真实反映神经表征,还是主要由语言模型重构,仍不明确。这一模糊性限制了可解释性,并阻碍了内在脑-语言对应关系的研究。为此,我们提出MD-SigLIP:一种边缘正则化的结构化语义对齐框架,直接在共享语义空间中对齐大脑嵌入与文本嵌入,实现基于检索的解码。该框架显式建模神经表征与语言语义间的对应关系。基于去重感知的Sigmoid对比学习,引入列表级边缘正则化项,强制正向语义簇与负样本间保持结构化排序。通过同时建模多正例语义结构与基于边缘的排序,该方法捕捉了语言嵌入在神经信号中体现的流形组织。实验表明,在全词汇与子集评估设置下均达到当前最优检索性能。
原文摘要 · Abstract (English)
With the rapid advancement of large language models, brain-language decoding has achieved remarkable progress. However, it remains unclear whether decoded content genuinely reflects neural representations or is largely reconstructed by the language model itself. This ambiguity limits interpretability and hinders the investigation of intrinsic brain-language correspondence. To address this challenge, we propose MD-SigLIP. This margin-regularized structured semantic alignment framework directly aligns brain embeddings with text embeddings in a shared semantic space, enabling retrieval-based decoding. This formulation enables explicit modeling of the correspondence between neural representations and language semantics. Building upon duplicate-aware sigmoid contrastive learning, we introduce a listwise margin-regularized term that enforces structured ranking constraints between positive semantic clusters and negative samples. By modeling multi-positive semantic structure and margin-based ordering simultaneously, the method captures the manifold organization of language embeddings reflected in neural signals. Experiments demonstrate state-of-the-art retrieval performance under both full-vocabulary and subset evaluation settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。