用聚合标签可提升模型鲁棒性,即使标签有噪声也不易失效。
Some Robustness Properties of Label Cleaning
- 通过聚合噪声标签信息,替代原始标签进行学习
- 在损失函数轻微错误时仍能收敛到最优分类器
- 适合处理标签不准确或数据收集有偏差的场景
我们证明,依赖聚合标签的学习方法(如从噪声响应中提炼标签信息)具有无法通过原始标签实现的鲁棒性。在风险一致性背景下,使用聚合标签的方法比使用原始标签能获得更强的一致性保证,即使目标损失(如零一分类误差)被近似为代理损失(通常为凸损失)。尽管经典统计假设下应充分利用所有信息(包括标签不确定性),但标准方法一旦损失函数存在轻微误设便失效;而利用聚合信息的方法仍能收敛至最优分类器。这表明,将数据收集、建模到预测全过程纳入考量,通过提炼噪声信号,可构建更稳健的分析方法。
原文摘要 · Abstract (English)
We demonstrate that learning procedures that rely on aggregated labels, e.g., label information distilled from noisy responses, enjoy robustness properties impossible without data cleaning. This robustness appears in several ways. In the context of risk consistency -- when one takes the standard approach in machine learning of minimizing a surrogate (typically convex) loss in place of a desired task loss (such as the zero-one mis-classification error) -- procedures using label aggregation obtain stronger consistency guarantees than those even possible using raw labels. And while classical statistical scenarios of fitting perfectly-specified models suggest that incorporating all possible information -- modeling uncertainty in labels -- is statistically efficient, consistency fails for ``standard'' approaches as soon as a loss to be minimized is even slightly mis-specified. Yet procedures leveraging aggregated information still converge to optimal classifiers, highlighting how incorporating a fuller view of the data analysis pipeline, from collection to model-fitting to prediction time, can yield a more robust methodology by refining noisy signals.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。