arXiv:2608.00949cs.LGstat.CO2026-08

根据数据自动选择最优损失函数,提升SVM鲁棒性与适应性。

Data-Driven Pinball-Loss Selection for Vertically Distributed Elastic-Net SVMs

  • 通过学习多个损失函数的加权组合,动态调整模型对误差的敏感度。
  • 理论证明:最优解性能不劣于任意固定参数的基线方法。
  • 适用于高维数据,支持分布式训练且结果与集中式一致。

传统的分位数损失支持向量机虽具鲁棒性,但其不对称参数常需手动设定。本文提出一种数据驱动的弹性网支持向量机,通过在候选分位数损失上学习单纯形约束权重,实现单个分类器的自适应损失融合。加权损失等价于一个依赖数据的有效分位数参数。经验奥拉克不等式表明:当权重正则化与单纯形截断趋近于零时,全局最优解的目标函数值不超过最佳固定候选方案;否则,超出部分可显式界定。针对高维数据,设计列分区变量分裂求解器,收敛速度达$O(1/T)$平方步长残差率。在标准初始化与全局参数下,任意列划分在精确算术中均产生与集中训练相同的迭代序列和最终解。实验验证了预测性能、数值等价性及多进程可扩展性。

原文摘要 · Abstract (English)

The pinball-loss support vector machine is robust, but its asymmetry parameter is usually fixed in advance. We propose a data-driven elastic-net support vector machine that learns simplex-constrained weights over candidate pinball losses while retaining one classifier. The weighted loss is equivalent to a pinball loss with a data-dependent effective parameter. An empirical oracle inequality shows that, when weight regularization and simplex truncation vanish, the classifier objective at a global minimizer does not exceed that of the best fixed candidate; otherwise, the excess is explicitly bounded. For high-dimensional data, we develop a column-partitioned variable-splitting solver. It converges with a best-iterate $O(1/T)$ squared-step residual rate. Under common initialization and global parameters, any column partition produces, in exact arithmetic, the same iterates and solution as centralized training. Experiments assess predictive behavior, numerical equivalence, and multi-process scalability.

SVM损失函数分布式学习弹性网

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。