arXiv:2601.08618stat.MLcs.LG2026-01

针对多标签二分类问题,提出基于配对AUC的低秩建模方法,提升预测鲁棒性。

Robust low-rank estimation with multiple binary responses using pairwise AUC loss

  • 通过最小化配对AUC代理损失,直接优化排序性能
  • 在高维与类别不平衡下仍保持良好表现,优于传统似然方法
  • 适用于需要稳定排序结果的场景,如医疗诊断与推荐系统

多二分类响应在现代数据分析中普遍存在。虽然对每个响应分别拟合逻辑回归计算高效,但忽略了任务间的共享结构,在高维和类别不平衡情况下统计效率低下。低秩模型可自然编码任务间的潜在依赖,但现有二值数据方法多基于似然且关注点分类而非排序性能。本文提出一种统一框架,直接通过最小化受试者工作特征曲线下面积(AUC)的代理损失来优化判别能力。该方法在多个响应间聚合配对AUC代理损失,并对系数矩阵施加低秩约束以利用共享结构。我们开发了一种基于截断奇异值分解的可扩展投影梯度下降算法。利用配对损失仅依赖线性预测值之差的特点,简化了计算与分析。建立了非渐近收敛保证,表明在适当正则条件下,收敛速度达到最优统计精度并呈线性收敛。大量模拟实验显示,该方法在标签切换与数据污染等挑战性场景下具有鲁棒性,且始终优于基于似然的方法。

原文摘要 · Abstract (English)

Multiple binary responses arise in many modern data-analytic problems. Although fitting separate logistic regressions for each response is computationally attractive, it ignores shared structure and can be statistically inefficient, especially in high-dimensional and class-imbalanced regimes. Low-rank models offer a natural way to encode latent dependence across tasks, but existing methods for binary data are largely likelihood-based and focus on pointwise classification rather than ranking performance. In this work, we propose a unified framework for learning with multiple binary responses that directly targets discrimination by minimizing a surrogate loss for the area under the ROC curve (AUC). The method aggregates pairwise AUC surrogate losses across responses while imposing a low-rank constraint on the coefficient matrix to exploit shared structure. We develop a scalable projected gradient descent algorithm based on truncated singular value decomposition. Exploiting the fact that the pairwise loss depends only on differences of linear predictors, we simplify computation and analysis. We establish non-asymptotic convergence guarantees, showing that under suitable regularity conditions, leading to linear convergence up to the minimax-optimal statistical precision. Extensive simulation studies demonstrate that the proposed method is robust in challenging settings such as label switching and data contamination and consistently outperforms likelihood-based approaches.

多任务学习低秩模型AUC优化鲁棒估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。