arXiv:2604.18827q-bio.NCcs.AI2026-04被引 8

用1500亿神经数据训练多模态脑模型,实现预测、解码与预报的灵活切换。

OmniMouse: Scaling properties of multi-modal, multi-task Brain Models on 150B Neural Tokens

论文配图:OmniMouse: Scaling properties of multi-modal, multi-task Brain Models on 150B Neural Tokens
图 1 · 摘自论文原文
  • 基于多任务多模态架构,统一处理神经信号、行为与刺激数据。
  • 在几乎所有评估场景中超越专用模型,但模型规模增益逐渐饱和。
  • 揭示脑建模中数据是瓶颈,未来大数据或催生质变新能力。

通过73只小鼠、323次实验采集的310万神经元数据(超1500亿神经令牌),我们训练了支持神经预测、行为解码和神经预报三种模式的多模态多任务模型。OmniMouse在多数测试场景中表现优于专用基线模型,验证了数据量增长带来稳定性能提升,但模型规模扩大带来的增益趋于饱和。这一反常现象表明:尽管语言与视觉领域依赖参数扩展,脑建模仍受数据限制,即便在小鼠视觉皮层这种相对简单系统中亦然。系统性缩放规律暗示神经建模可能经历相变,更大更丰富的数据集或触发类大语言模型的涌现特性。代码开源于https://github.com/enigma-brain/omnimouse。

原文摘要 · Abstract (English)

Scaling data and artificial neural networks has transformed AI, driving breakthroughs in language and vision. Whether similar principles apply to modeling brain activity remains unclear. Here we leveraged a dataset of 3.1 million neurons from the visual cortex of 73 mice across 323 sessions, totaling more than 150 billion neural tokens recorded during natural movies, images and parametric stimuli, and behavior. We train multi-modal, multi-task models that support three regimes flexibly at test time: neural prediction, behavioral decoding, neural forecasting, or any combination of the three. OmniMouse achieves state-of-the-art performance, outperforming specialized baselines across nearly all evaluation regimes. We find that performance scales reliably with more data, but gains from increasing model size saturate. This inverts the standard AI scaling story: in language and computer vision, massive datasets make parameter scaling the primary driver of progress, whereas in brain modeling -- even in the mouse visual cortex, a relatively simple system -- models remain data-limited despite vast recordings. The observation of systematic scaling raises the possibility of phase transitions in neural modeling, where larger and richer datasets might unlock qualitatively new capabilities, paralleling the emergent properties seen in large language models. Code available at https://github.com/enigma-brain/omnimouse.

脑建模多模态数据驱动神经科学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。