用神经机制启发的模型,让AI更快更省地预选信息。
Scalable Machines with Intrinsic Higher Mental-State Dynamics
- 通过三元调制环路预选关键信息,模拟清醒想象的计算原理。
- 在ImageNet-1K上训练速度更快,所需注意力头、层数和令牌更少。
- 适合追求高效推理与可扩展性的视觉与语言模型研究者。
基于新近的细胞神经生物学突破及将新皮层锥体神经元与不同精神状态关联的生物物理建模,本文提出一种数学严谨的框架,表明模型(如Transformer)可通过查询(Q)、键(K)、值(V)间的三元调制回路,在注意力应用前实现与清醒想象力相关的计算原则,以预选相关信息。在ImageNet-1K上的可扩展性实验表明,相比标准视觉Transformer(ViT),该方法显著加快学习速度并降低计算开销(减少注意力头、层数和输入令牌数),其复杂度约为输入令牌数N的$/mathcal{O}(N)$阶,与此前在强化学习和语言建模中的发现一致。
原文摘要 · Abstract (English)
Drawing on recent breakthroughs in cellular neurobiology and detailed biophysical modeling linking neocortical pyramidal neurons to distinct mental-state regimes, this work introduces a mathematically grounded formulation showing how models (e.g., Transformers) can implement computational principles underlying awake imaginative thought to pre-select relevant information before attention is applied via triadic modulation loops among queries ($Q$), keys ($K$), and values ($V$).~Scalability experiments on ImageNet-1K, benchmarked against a standard Vision Transformer (ViT), demonstrate significantly faster learning with reduced computational demand (fewer heads, layers, and tokens), consistent with our prior findings in reinforcement learning and language modeling. The approach operates at approximately $\mathcal{O}(N)$ complexity with respect to the number of input tokens $N$.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。