首次研究两阶段学习拒决系统的抗攻击能力,提升复杂环境下的决策可靠性。
Adversarial Robustness in Two-Stage Learning-to-Defer: Algorithms and Guarantees
- 提出无目标与有目标两类攻击,破坏最优任务分配
- 设计SARD算法,在多种任务中保持鲁棒性与准确率
- 适用于需要安全可靠的多代理决策场景
两阶段学习拒决(L2D)通过将输入分配给主模型或多个离线专家,实现最优任务委派,支持复杂多代理环境中的可靠决策。然而,现有L2D框架假设输入干净,易受对抗扰动影响,导致错误调度或专家过载。本文首次系统研究两阶段L2D的对抗鲁棒性,提出两种新型攻击策略——无目标攻击扰乱最优分配,有目标攻击强制查询指向特定代理。为防御此类威胁,我们提出SARD,一种基于一组可证明贝叶斯一致和$(/mathcal{R}, /mathcal{G})$一致代理损失的凸学习算法,该保证在分类、回归及多任务设置下均成立。实验表明,SARD在对抗攻击下显著提升鲁棒性,同时保持优异的原始性能,是迈向安全可信L2D部署的关键一步。
原文摘要 · Abstract (English)
Two-stage Learning-to-Defer (L2D) enables optimal task delegation by assigning each input to either a fixed main model or one of several offline experts, supporting reliable decision-making in complex, multi-agent environments. However, existing L2D frameworks assume clean inputs and are vulnerable to adversarial perturbations that can manipulate query allocation--causing costly misrouting or expert overload. We present the first comprehensive study of adversarial robustness in two-stage L2D systems. We introduce two novel attack strategie--untargeted and targeted--which respectively disrupt optimal allocations or force queries to specific agents. To defend against such threats, we propose SARD, a convex learning algorithm built on a family of surrogate losses that are provably Bayes-consistent and $(\mathcal{R}, \mathcal{G})$-consistent. These guarantees hold across classification, regression, and multi-task settings. Empirical results demonstrate that SARD significantly improves robustness under adversarial attacks while maintaining strong clean performance, marking a critical step toward secure and trustworthy L2D deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。