arXiv:2506.06840stat.MLcs.AI2025-06被引 2

为LSTM模型提供一套系统化选择方法,降低调参成本

A Statistical Framework for Model Selection in LSTM Networks

  • 基于信息准则与收缩估计,构建适配时序结构的惩罚似然函数
  • 提出广义阈值法处理隐藏状态动态,提升建模灵活性
  • 适用于生物医学时序数据,适合需要稳定性能的工程场景

长短期记忆(LSTM)神经网络在自然语言处理、时间序列预测等众多领域已成为时序建模的核心工具。尽管成果显著,模型选择问题——包括超参数调优、架构设计与正则化选择——仍主要依赖经验性方法且计算成本高。本文提出一种统一的统计框架,用于系统性地进行LSTM模型选择。该框架将经典模型选择思想(如信息准则与收缩估计)拓展至序列神经网络,定义了适应时序结构的惩罚似然函数,提出针对隐藏状态动态的广义阈值方法,并采用变分贝叶斯和近似边缘似然方法实现高效估计。多个以生物医学数据为中心的实例验证了该框架的灵活性与性能提升。

原文摘要 · Abstract (English)

Long Short-Term Memory (LSTM) neural network models have become the cornerstone for sequential data modeling in numerous applications, ranging from natural language processing to time series forecasting. Despite their success, the problem of model selection, including hyperparameter tuning, architecture specification, and regularization choice remains largely heuristic and computationally expensive. In this paper, we propose a unified statistical framework for systematic model selection in LSTM networks. Our framework extends classical model selection ideas, such as information criteria and shrinkage estimation, to sequential neural networks. We define penalized likelihoods adapted to temporal structures, propose a generalized threshold approach for hidden state dynamics, and provide efficient estimation strategies using variational Bayes and approximate marginal likelihood methods. Several biomedical data centric examples demonstrate the flexibility and improved performance of the proposed framework.

LSTM模型选择统计学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。