arXiv:2508.01575cs.LG2025-08被引 2

用可学习的非线性函数提升长期时间序列预测精度。

KANMixer: a minimal KAN-centered mixer for long-term time series forecasting

  • 以KAN为核心构建极简架构,通过自适应基函数建模复杂非线性。
  • 28个基准场景中16次最低均方误差,11次最低平均绝对误差。
  • 发现B样条基优于傅里叶与小波,且深度适中效果最佳。

长期时间序列预测(LTSF)支撑能源管理、天气预报等关键应用,但多步预测的准确性仍难保证。现有方法以MLP和Transformer为主,依赖简单线性映射或复杂手工归纳偏置,引发疑问:是否可用更表达性强、原则性更强的非线性核心替代?为此,我们探究了柯尔莫戈洛夫-阿诺德网络(KAN)在LTSF中的潜力,其特点为可自适应调整的基函数,能精细调控非线性。提出KANMixer,一种极简的KAN中心架构,包含多尺度池化前端、基于KAN的时序混合主干和预测头。通过避免冗余模块,清晰评估KAN组件的作用。在28个基准时序设置下对比9种基线,KANMixer在16个场景中实现最低均方误差(MSE),在11个场景中实现最低平均绝对误差(MAE)。三组代表性数据集的消融实验表明:KAN性能高度依赖边函数选择;B样条基优于傅里叶和小波;预测头贡献最大;中等深度优于更深且不稳定的堆叠;分解先验对MLP有益,却损害KAN。这些结果不仅为集成KAN提供实用指导,更揭示结构先验与骨干非线性的深层依赖关系:对MLP有益的设计可能削弱KAN表现。

原文摘要 · Abstract (English)

Long-term time series forecasting (LTSF) underpins critical applications from energy management to weather prediction, yet achieving reliable multi-step-ahead accuracy remains challenging. Existing LTSF approaches, dominated by MLP- and Transformer-based architectures, either rely on simple linear mappings or introduce increasingly complex hand-crafted inductive biases, raising the question of whether a more expressive and principled nonlinear core could offer a better alternative. Therefore, we investigate whether Kolmogorov-Arnold Networks (KANs), a recently proposed model featuring adaptive basis functions capable of granular modulation of nonlinearities, can improve LTSF performance, and under which design choices they are most effective. Specifically, we propose KANMixer, a minimal KAN-centered architecture consisting of a multi-scale pooling frontend, a KAN-based temporal mixing backbone, and prediction heads. By avoiding heavy auxiliary modules, KANMixer enables a clear assessment of KAN components in LTSF. Across 28 benchmark-horizon settings against nine baselines, KANMixer achieves the best MSE in 16 settings and the best MAE in 11. Furthermore, extensive ablations on three representative datasets show that KAN effectiveness depends strongly on the choice of edge function; B-spline bases outperform Fourier and Wavelet alternatives; the prediction head contributes most to the gains; moderate depth is preferred over deeper unstable stacks; and decomposition priors help MLP but harm KAN. Beyond practical guidance for integrating KAN into LTSF, these results reveal an underexplored dependency between structural priors and backbone nonlinearity: design choices that benefit MLP can degrade KAN.

时间序列KAN非线性建模预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。