arXiv:2608.13601cs.LG2026-08

探究不确定性采样在噪声标签下的失效原因,发现错误位置影响有限,效果依赖数据集和评估指标。

Hard Cases, Bad Labels: Testing Error Exposure and Error Location in Uncertainty Sampling Under Bounded Label Noise

  • 对比不确定性与随机采样,在三种噪声下测试其表现差异。
  • 在乳腺癌数据集上,难度相关噪声比随机噪声更损害模型性能,但其他数据集无此趋势。
  • 结果表明采样鲁棒性受数据集、预算和评价指标共同影响,非普遍适用。

主动学习通过选择信息量大的样本降低标注成本,但最不确定的样本往往最难正确标注。本研究检验不确定性采样是否因获取更多错误标签,或因困难区域的误差集中而失效。在三个公开二分类表格数据集上,对比基于边界(margin-based)的不确定性采样与随机采样,在干净标签、随机分类噪声(RCN)及有界难度依赖噪声条件下进行实验。实验设置100组配对种子、9种预期噪声率(0至0.30)、标注预算20至120,并在每个预算下通过交叉验证重新选择带正则化的逻辑回归模型。使用暴露匹配的RCN对照组控制最终获取的错误率均值,清洁标签扩展实验达预算400。在干净标签下,不确定性采样使所有数据集的归一化平衡准确率学习曲线下面积提升1.09至1.77个百分点。在乳腺癌威斯康星数据集上,难度依赖噪声在六种噪声率下比RCN造成更大损失,但在银行票据认证和MAGIC伽马射线望远镜数据集上未观察到类似现象。暴露匹配分析未发现结构化误差位置存在普遍额外惩罚。在清洁版MAGIC数据上,不确定性采样虽提升平衡准确率,但固定假阳性率下平均精度和真正例率下降。因此,不确定性采样具标签效率,但其表面鲁棒性取决于数据集、预算、噪声结构和评估指标。

原文摘要 · Abstract (English)

Active learning can reduce labeling cost by selecting informative examples, but the most uncertain examples may also be the hardest to label correctly. This study tests whether uncertainty sampling fails because it acquires more corrupted labels or because errors concentrated in difficult regions are especially harmful. Margin-based uncertainty sampling is compared with random sampling under clean labels, random classification noise (RCN), and bounded difficulty-dependent noise on three public binary tabular datasets. The design uses 100 paired seeds, nine expected noise rates from 0 to 0.30, annotation budgets from 20 to 120, and logistic regression with regularization re-selected by cross-validation at every budget. An exposure-matched RCN control aligns mean final acquired corruption, while a clean-label extension reaches budget 400. Under clean labels, uncertainty sampling improved normalized balanced-accuracy area under the learning curve by 1.09 to 1.77 percentage points on all datasets. Difficulty-dependent noise reduced this advantage more than RCN at six of eight rates on Breast Cancer Wisconsin, but at no tested rate on Banknote Authentication or MAGIC Gamma Telescope. Exposure-matched analyses found no corrected evidence for a universal additional penalty from structured error location. On clean MAGIC data, uncertainty sampling improved balanced accuracy while reducing average precision and true-positive rate at fixed false-positive rates. Thus, uncertainty sampling was label-efficient, but its apparent robustness depended on dataset, budget, noise structure, and evaluation metric.

主动学习不确定性采样噪声标签实验分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。