arXiv:2409.09674cs.LGstat.ML2024-09

无需先验信息,用排序法选出最优简约模型。

Model Selection Through Model Sorting

  • 基于嵌套模型的可区分性,构造连续经验过风险边界。
  • 提出S-NER方法,排序后模型最小风险显著下降。
  • 适合无先验知识的模型选择,尤其在回归与分类中表现优。

我们提出一种新方法来选择最佳数据模型。基于嵌套模型的独有性质,找到包含风险最小预测器的最简约模型。证明了两个连续嵌套模型的最小经验风险差存在可能近似正确(PAC)界,称为连续经验过风险(SEER)。基于此界,提出嵌套经验风险(NER)模型选择方法。通过排序NER(S-NER)智能排序模型,使最小风险递减。构建检验以预测扩展模型是否降低最小风险。高概率下,NER选择真实模型阶数,S-NER选择包含风险最小预测器的最简约模型。在线性回归中应用S-NER,发现其无需任何先验信息即可超越借助真实模型阶数先验的正交匹配追踪(OMP)算法的精度。在UCR数据集上,NER方法大幅降低分类复杂度,仅损失微小准确率。

原文摘要 · Abstract (English)

We propose a novel approach to select the best model of the data. Based on the exclusive properties of the nested models, we find the most parsimonious model containing the risk minimizer predictor. We prove the existence of probable approximately correct (PAC) bounds on the difference of the minimum empirical risk of two successive nested models, called successive empirical excess risk (SEER). Based on these bounds, we propose a model order selection method called nested empirical risk (NER). By the sorted NER (S-NER) method to sort the models intelligently, the minimum risk decreases. We construct a test that predicts whether expanding the model decreases the minimum risk or not. With a high probability, the NER and S-NER choose the true model order and the most parsimonious model containing the risk minimizer predictor, respectively. We use S-NER model selection in the linear regression and show that, the S-NER method without any prior information can outperform the accuracy of feature sorting algorithms like orthogonal matching pursuit (OMP) that aided with prior knowledge of the true model order. Also, in the UCR data set, the NER method reduces the complexity of the classification of UCR datasets dramatically, with a negligible loss of accuracy.

模型选择嵌套模型线性回归特征排序

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。