arXiv:2504.16585cs.LGstat.CO2025-04

利用标注噪声提升逻辑回归变量选择效果

Leveraging Noisy Manual Labels as Useful Information: An Information Fusion Approach for Enhanced Variable Selection in Penalized Logistic Regression

  • 将人工标注噪声视为信息源,融合进正则化逻辑回归的变量选择
  • 在多个大规模数据集上,变量估计和分类性能均优于传统方法
  • 适合处理分布式大数据且对数据划分不敏感,适合工业级应用

在大规模监督学习中,正则化逻辑回归(PLR)通过正则化有效缓解过拟合,但其性能高度依赖于稳健的变量选择。本文表明,人工标注过程中引入的标签噪声,常被视为瑕疵,实则可作为有价值的信息源,用于提升PLR中的变量选择效果。理论上证明,此类噪声与分类难度密切相关,能比仅使用真实标签更精确地估计非零系数,从而将常见缺陷转化为有用信息资源。为在无法集中存储数据的大规模场景中高效实现该信息融合,我们提出一种基于交替方向乘子法(ADMM)的新型分区无关并行算法。该方法确保解在数据跨工作节点分布时保持不变,是实现可复现、稳定分布式学习的关键特性,并保证全局收敛速度达到次线性。在多个大规模数据集上的大量实验表明,所提方法在变量估计精度和分类性能上持续优于传统变量选择技术,证实了有意识融合噪声标签的显著价值。

原文摘要 · Abstract (English)

In large-scale supervised learning, penalized logistic regression (PLR) effectively mitigates overfitting through regularization, yet its performance critically depends on robust variable selection. This paper demonstrates that label noise introduced during manual annotation, often dismissed as a mere artifact, can serve as a valuable source of information to enhance variable selection in PLR. We theoretically show that such noise, intrinsically linked to classification difficulty, helps refine the estimation of non-zero coefficients compared to using only ground truth labels, effectively turning a common imperfection into a useful information resource. To efficiently leverage this form of information fusion in large-scale settings where data cannot be stored on a single machine, we propose a novel partition insensitive parallel algorithm based on the alternating direction method of multipliers (ADMM). Our method ensures that the solution remains invariant to how data is distributed across workers, a key property for reproducible and stable distributed learning, while guaranteeing global convergence at a sublinear rate. Extensive experiments on multiple large-scale datasets show that the proposed approach consistently outperforms conventional variable selection techniques in both estimation accuracy and classification performance, affirming the value of intentionally fusing noisy manual labels into the learning process.

变量选择正则化分布式学习标签噪声

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。