arXiv:2505.01130cs.LGcs.AI2025-05被引 3

提出一套无需额外数据的模型抗攻击评估框架,适用于多种攻击类型。

Risk Analysis and Design Against Adversarial Actions

  • 基于松弛优化构建通用评估框架,适配支持向量回归等模型。
  • 无需测试数据即可评估模型在部署时的脆弱性,支持分布无关设定。
  • 适合关注模型鲁棒性与可信度的研究者,尤其适用于对抗样本分析。

近年来,如何使学习模型在面对部署阶段的对抗性行为时仍能提供可靠预测,已成为机器学习领域的核心问题。这一挑战源于模型在实际部署中遇到的数据常偏离训练时的假设条件。本文针对部署阶段的对抗性行为,提出一种通用且理论严谨的框架,用于评估模型对不同类型的攻击及强度的鲁棒性。初始聚焦于支持向量回归(SVR),该方法可自然扩展至通过松弛优化进行学习的广泛领域。其结果可在无需额外测试数据的情况下评估模型脆弱性,且在分布无关设定下运行。这些成果不仅有助于提升对模型适用性的信任,还可辅助在多个候选模型间进行选择。此外,本文还表明,该框架能为分布外(out-of-distribution)场景下的新发现提供有益启示。

原文摘要 · Abstract (English)

Learning models capable of providing reliable predictions in the face of adversarial actions has become a central focus of the machine learning community in recent years. This challenge arises from observing that data encountered at deployment time often deviate from the conditions under which the model was trained. In this paper, we address deployment-time adversarial actions and propose a versatile, well-principled framework to evaluate the model's robustness against attacks of diverse types and intensities. While we initially focus on Support Vector Regression (SVR), the proposed approach extends naturally to the broad domain of learning via relaxed optimization techniques. Our results enable an assessment of the model vulnerability without requiring additional test data and operate in a distribution-free setup. These results not only provide a tool to enhance trust in the model's applicability but also aid in selecting among competing alternatives. Later in the paper, we show that our findings also offer useful insights for establishing new results within the out-of-distribution framework.

对抗攻击模型鲁棒性风险评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。