arXiv:2601.02193cs.LGcs.DS2026-01被引 5

研究数据污染对学习算法的影响,发现看似有益的篡改会破坏最优算法性能。

Learning with Monotone Adversarial Corruptions

  • 引入单调对抗性污染模型,模拟恶意添加带标签数据点
  • 现有最优分类算法在新测试点上误差显著上升
  • 基于一致收敛的算法仍保持稳定,适合对鲁棒性要求高场景

我们通过引入单调对抗性污染模型,研究标准机器学习算法对数据交换性和独立性的依赖程度。在此模型中,攻击者观察一个干净的i.i.d.数据集后,可插入任意数量的被污染点,这些点需满足单调性约束:其标签由真实目标函数决定。出人意料的是,我们证明所有已知的二分类最优学习算法均可能在来自相同分布的新测试点上产生次优期望误差。相反,基于一致收敛的算法其保证不降。结果揭示了最优学习算法在面对看似有益的单调污染时的脆弱性,暴露了其对交换性的过度依赖。

原文摘要 · Abstract (English)

We study the extent to which standard machine learning algorithms rely on exchangeability and independence of data by introducing a monotone adversarial corruption model. In this model, an adversary, upon looking at a "clean" i.i.d. dataset, inserts additional "corrupted" points of their choice into the dataset. These added points are constrained to be monotone corruptions, in that they get labeled according to the ground-truth target function. Perhaps surprisingly, we demonstrate that in this setting, all known optimal learning algorithms for binary classification can be made to achieve suboptimal expected error on a new independent test point drawn from the same distribution as the clean dataset. On the other hand, we show that uniform convergence-based algorithms do not degrade in their guarantees. Our results showcase how optimal learning algorithms break down in the face of seemingly helpful monotone corruptions, exposing their overreliance on exchangeability.

对抗训练泛化能力数据污染

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。