arXiv:2511.04979cs.LGstat.CO2025-11中稿 · Stat被引 1

提出高效ROC-SVM方法,显著降低训练时间同时保持高分类性能。

Scaling Up ROC-Optimizing Support Vector Machines

  • 用不完全U统计量替代全对比较,降低计算复杂度
  • 在真实与合成数据集上达到与原版相当的AUC表现
  • 适合处理大规模不平衡分类问题的研究者

ROC-SVM由Rakotomamonjy首次提出,直接最大化受试者工作特征曲线下面积(AUC),在类别不平衡场景下成为传统二分类的有效替代。然而其实际应用受限于高昂的计算成本,因训练需评估所有O(n²)样本对。为此,本文提出一种可扩展的ROC-SVM变体,利用不完全U统计量显著降低计算复杂度。进一步通过低秩核近似将框架拓展至非线性分类,实现在再生核希尔伯特空间中的高效训练。理论分析建立了误差界以证明近似的合理性,实验证明该方法在合成与真实数据集上均能实现与原始ROC-SVM相当的AUC性能,且训练时间大幅缩减。

原文摘要 · Abstract (English)

The ROC-SVM, originally proposed by Rakotomamonjy, directly maximizes the area under the ROC curve (AUC) and has become an attractive alternative of the conventional binary classification under the presence of class imbalance. However, its practical use is limited by high computational cost, as training involves evaluating all $O(n^2)$. To overcome this limitation, we develop a scalable variant of the ROC-SVM that leverages incomplete U-statistics, thereby substantially reducing computational complexity. We further extend the framework to nonlinear classification through a low-rank kernel approximation, enabling efficient training in reproducing kernel Hilbert spaces. Theoretical analysis establishes an error bound that justifies the proposed approximation, and empirical results on both synthetic and real datasets demonstrate that the proposed method achieves comparable AUC performance to the original ROC-SVM with drastically reduced training time.

SVMAUC优化大规模学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。