用6千参数实现高速高精度序列预测,可解释性强。
ALPHABET: A Laplace-Pole History Aggregator with Banked Exponential Transport

- 用复数极点模式压缩历史信息,分两路独立处理并读取能量与滞后量。
- 在82个任务上平均排名3.97,推理速度比基线快5.02倍。
- 模型小巧可审计,适合对效率与可解释性要求高的场景。
能否仅用数千参数和可审计的预测接口保持序列模型竞争力?我们提出ALPHABET,一种线性时间紧凑模型,将时间历史压缩为稳定的复数极点模式:一个直接银行将模态状态合成回特征轨迹,一个独立级联银行分析变换后的轨迹而不重合成,最后的仿射头仅读取两个银行的模态能量与滞后矩。我们刻画了该描述符保留的时间信息:对于平稳、完全观测的特征过程,每个模态能量是二阶谱的频域局部测量,这些测量的连续统可确定谱,几乎每个模态能分离任意固定有限的谱区分类别。在匹配低滞后统计的高斯控制下,学习到的描述符逼近贝叶斯最优,而原始自协方差仍处于随机水平。在固定的82任务注册表中,ALPHABET在十家族完整对比中取得平均排名3.97。在常见宽度D=64的运行基准下,其6,437个参数实现了平均5.02倍更快的推理和3.93倍更快的完整训练步数,优于九个基线。
原文摘要 · Abstract (English)
Can a sequence model remain competitive with only a few thousand parameters and an explicitly auditable prediction interface? We introduce ALPHABET, a compact linear-time model that compresses temporal history into stable complex pole modes: a direct bank synthesizes its modal states back into the feature trajectory, an independent cascaded bank analyzes the transformed trajectory without resynthesis, and an affine head reads only modal energies and lag moments from both banks. We characterize the temporal information this descriptor retains: for a stationary, fully observed feature process, each mode energy is a frequency-localized measurement of the second-order spectrum, the continuum of such measurements identifies the spectrum, and almost every mode separates any fixed finite set of spectrally distinct classes. On a Gaussian control with matched low-lag statistics, the learned descriptor approaches the Bayes oracle where raw autocovariances remain at chance. Across the fixed 82-task registry, ALPHABET attains mean rank 3.97 in the complete ten-family comparison. At the common-width D=64 runtime anchor, its 6,437 parameters deliver 5.02 times faster inference and 3.93 times faster complete training steps than the nine baselines on average.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。