首个融合脑科学、视觉与语言的统一模型,实现双向生成与理解。
BrainJanus: A Unified Model for Understanding and Generation across Brain, Vision, and Language

- 用统一令牌化将脑电活动转为与视觉语言对齐的离散符号
- 支持任意模态间生成,如脑信号生成图像或文字,性能领先
- 零样本泛化且保持生物拓扑可解释性,适合跨模态神经科学研究
建模外部感官刺激与内部神经活动之间的双向对应关系,已成为神经科学的关键前沿。然而,现有方法多将脑编码与解码视为孤立任务,依赖单模态对齐和外部先验,忽视了大脑作为多模态整合系统的内在特性。为此,我们提出 BrainJanus,首个将脑科学、视觉与语言统一于单一框架的模型。具体地,我们引入统一脑令牌化器,将连续神经动态量化为与视觉和语言表示对齐的离散令牌,位于共享的全模态空间中。基于此,我们采用端到端自回归架构,通过预测下一个令牌实现任意模态间的无缝生成,涵盖图像→脑、文本→脑编码,以及脑→图像、脑→文本解码。大量实验表明,BrainJanus 在多个基准上均取得优异性能。此外,该框架展现出零样本泛化能力,并保留可解释的生物学拓扑结构,凸显其作为通用脑建模范式潜力。代码已开源。
原文摘要 · Abstract (English)
Modeling the bidirectional correspondence between external sensory stimuli and internal neural activity has emerged as a critical frontier in neuroscience. However, existing approaches predominantly treat brain encoding and decoding as isolated tasks, relying heavily on unimodal alignment and external priors while overlooking the brain's intrinsic nature as a multimodal integration system. To address these limitations, we propose BrainJanus, the first unified brain model that integrates brain, vision, and language within a single framework. Specifically, we introduce a Unified Brain Tokenizer to quantize continuous neural dynamics into discrete tokens aligned with visual and linguistic representations in a shared Omni space. Building on this, we utilize an All-in-One autoregressive architecture that leverages next-token prediction to enable seamless any-to-any generation, which encompasses image-to-brain and text-to-brain encoding, and brain-to-image and brain-to-text decoding. Extensive experiments demonstrate that BrainJanus achieves superior performance across diverse benchmarks. Furthermore, our framework exhibits zero-shot generalization and preserves interpretable biological topography, highlighting its potential as a general-purpose brain modeling paradigm. The code is available at \href{https://github.com/HaitaoWuTJU/BrainJanus}{GitHub}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。