首个针对单阶段学习拒决的抗扰动框架,提升模型在对抗攻击下的可靠性。
Adversarial Robustness in One-Stage Learning-to-Defer
- 提出单阶段学习拒决的对抗攻击建模与代价敏感损失函数
- 在多个基准数据集上实现对抗攻击下性能保持与拒决决策鲁棒性提升
- 适用于需可靠决策的医疗、金融等高风险场景
学习拒决(L2D)通过将输入分配给预测器或外部专家,实现混合决策。尽管前景广阔,但现有方法极易受到对抗扰动影响,扰动不仅会改变预测结果,还会操纵拒决决策。以往的鲁棒性分析局限于两阶段设置,而未涵盖预测器与分配策略联合训练的一阶段端到端情形。本文首次提出一阶段学习拒决的对抗鲁棒性框架,涵盖分类与回归任务。方法包括攻击形式化、代价敏感对抗代理损失设计,并建立理论保证,如$θ$、$(ρ, φ)$及贝叶斯一致性。在多个基准数据集上的实验表明,该方法在无目标与有目标攻击下均显著提升鲁棒性,同时保持原始干净数据性能。
原文摘要 · Abstract (English)
Learning-to-Defer (L2D) enables hybrid decision-making by routing inputs either to a predictor or to external experts. While promising, L2D is highly vulnerable to adversarial perturbations, which can not only flip predictions but also manipulate deferral decisions. Prior robustness analyses focus solely on two-stage settings, leaving open the end-to-end (one-stage) case where predictor and allocation are trained jointly. We introduce the first framework for adversarial robustness in one-stage L2D, covering both classification and regression. Our approach formalizes attacks, proposes cost-sensitive adversarial surrogate losses, and establishes theoretical guarantees including $\mathcal{H}$, $(\mathcal{R }, \mathcal{F})$, and Bayes consistency. Experiments on benchmark datasets confirm that our methods improve robustness against untargeted and targeted attacks while preserving clean performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。