arXiv:2605.12992q-bio.NCcs.LG2026-05

首个大规模神经元放电预测基准,揭示了传统评估的局限性。

SpikeProphecy: A Large-Scale Benchmark for Autoregressive Neural Population Forecasting

论文配图:SpikeProphecy: A Large-Scale Benchmark for Autoregressive Neural Population Forecasting
图 1 · 摘自论文原文
  • 提出分解指标,分离时间精度、空间模式和幅度无关对齐三个维度。
  • 在近9万神经元数据上验证,不同脑区预测能力存在稳定排序。
  • 发现子泊松评估下存在生物物理限制,且模型蒸馏无效。

神经种群模型通常通过预测与实际脉冲计数间的皮尔逊相关系数 $r$ 进行评估,这一单一数值掩盖了关键结构。我们主张评估方式与模型构建同等重要,提出 SpikeProphecy,首个基于真实电生理记录的因果自回归脉冲计数预测大规模基准。核心贡献是种群指标分解,将整体性能拆分为时间保真度、空间模式准确性和幅度不变对齐。该分解揭示了聚合标量所掩盖的数据内在结构。我们在105个 Neuropixels 会话(Steinmetz 2019 + IBL Repeated Site;约89,800个神经元)上应用此协议,涵盖七种架构基线:四种状态空间模型(三类对角、一类非对角)、一个Transformer、一个LSTM及一个脉冲网络。分解结果显示跨七种基线一致的脑区可预测性排序,且经协方差校正后仍显著高于发放统计约束(区域 $ΔR^2 = 0.018$)。同时揭示子泊松评估下存在真实生物物理约束的下限,并表明在泊松计数域中基于输出率的KL蒸馏在ANN到SNN迁移中无效。

原文摘要 · Abstract (English)

Neural population models, which predict the joint firing of many simultaneously recorded neurons forward in time, are typically evaluated by a single aggregate Pearson correlation $r$ between predicted and actual spike counts, a number that masks critical structure. We argue that how we evaluate spike forecasting matters as much as what we build, and introduce SpikeProphecy, the first large-scale benchmark for causal, autoregressive spike-count forecasting on real electrophysiology recordings. Our core contribution is a population metric decomposition that separates aggregate performance into temporal fidelity, spatial pattern accuracy, and magnitude-invariant alignment. The decomposition surfaces aspects of the underlying data that an aggregate scalar collapses together. We apply the protocol to 105 Neuropixels sessions (Steinmetz 2019 + IBL Repeated Site; ~89,800 neurons) with seven architecture baselines spanning four structural families: four SSMs (three diagonal and one non-diagonal), a Transformer, an LSTM, and a spiking network. The decomposition surfaces a brain-region predictability ranking that reproduces across all seven baselines and survives ANCOVA correction for firing-statistics constraints (region $ΔR^2 = 0.018$ above the firing-statistics covariates). It also exposes a sub-Poisson evaluation floor where rigorous metrics combine with genuine biophysical constraints on regular spike trains, and yields a negative result on KL-on-output-rates distillation for ANN-to-SNN transfer in this Poisson count domain.

神经建模脉冲预测评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。