研究如何在含恶意标签数据流中稳定学习,允许不确定时放弃预测。
Adversarial Resilience against Clean-Label Attacks in Realizable and Noisy Settings
- 设计可弃权的学习机制,对恶意标签样本免费放弃预测
- 首次在噪声环境下分析基于分歧的阈值学习器的鲁棒性
- 突破理想设定限制,拓展至标签随机的现实场景
我们研究在独立同分布数据流中顺序学习时,如何在包含未知数量干净标签对抗样本的情况下建立类似随机性的保证。允许学习者在不确定时放弃预测,其损失由误分类和弃权误差共同衡量,且对对抗样本的弃权行为不计代价。该方法基于Goel等(arXiv:2306.13119)的工作,我们修正了其论证中的不准确之处。然而原方法仅适用于可实现情形,即标签由假设空间内的某个函数 $f^*$ 决定。本文进一步基于类似思路,将方法扩展至泛化设定(agnostic setting),其中标签具有随机性。首次在泛化背景下引入干净标签对抗者概念,对受干净标签对抗与噪声影响的阈值学习器进行理论分析。
原文摘要 · Abstract (English)
We investigate the challenge of establishing stochastic-like guarantees when sequentially learning from a stream of i.i.d. data that includes an unknown quantity of clean-label adversarial samples. We permit the learner to abstain from making predictions when uncertain. The regret of the learner is measured in terms of misclassification and abstention error, where we allow the learner to abstain for free on adversarial injected samples. This approach is based on the work of Goel, Hanneke, Moran, and Shetty from arXiv:2306.13119. We explore the methods they present and manage to correct inaccuracies in their argumentation. However, this approach is limited to the realizable setting, where labels are assigned according to some function $f^*$ from the hypothesis space $\mathcal{F}$. Based on similar arguments, we explore methods to make adaptations for the agnostic setting where labels are random. Introducing the notion of a clean-label adversary in the agnostic context, we are the first to give a theoretical analysis of a disagreement-based learner for thresholds, subject to a clean-label adversary with noise.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。