arXiv:2412.07326cs.LG2024-12被引 2

提出新方法生成更隐蔽的表格数据攻击样本,提升检测准确性。

Addressing Key Challenges of Adversarial Attacks and Defenses in the Tabular Domain: A Methodological Framework for Coherence and Consistency

  • 通过保持特征依赖关系生成连贯的对抗样本
  • 设计类特定异常检测法,识别细微扰动
  • 结合SHAP分析模型决策不一致,适合安全研究者

基于表格数据的机器学习模型易受对抗攻击,尤其在仅能访问模型输出的现实场景中。由于表格数据特征间存在复杂依赖关系,对抗样本必须保持一致性以避免被识别。现有评估指标(成功率、扰动幅度、查询次数)未能充分反映这一挑战。为此,本文提出一种在保留特征依赖的前提下扰动数据的方法,并引入类特定异常检测(CSAD),从预测类别分布而非整体正常分布角度评估样本异常性,有效识别其他类别下看似合理但实际异常的微小扰动。结合SHAP可解释性技术,扩展出基于SHAP的异常检测机制,通过双重评估(异常检测率与SHAP分析)全面衡量对抗样本质量。在四个目标模型上对黑盒查询型和基于迁移的梯度攻击进行实验,使用基准表格数据集验证,揭示攻击者风险、投入与攻击质量间的差异与权衡,为未来表格数据领域的对抗攻防研究提供基础支持。

原文摘要 · Abstract (English)

Machine learning models trained on tabular data are vulnerable to adversarial attacks, even in realistic scenarios where attackers only have access to the model's outputs. Since tabular data contains complex interdependencies among features, it presents a unique challenge for adversarial samples which must maintain coherence and respect these interdependencies to remain indistinguishable from benign data. Moreover, existing attack evaluation metrics-such as the success rate, perturbation magnitude, and query count-fail to account for this challenge. To address those gaps, we propose a technique for perturbing dependent features while preserving sample coherence. In addition, we introduce Class-Specific Anomaly Detection (CSAD), an effective novel anomaly detection approach, along with concrete metrics for assessing the quality of tabular adversarial attacks. CSAD evaluates adversarial samples relative to their predicted class distribution, rather than a broad benign distribution. It ensures that subtle adversarial perturbations, which may appear coherent in other classes, are correctly identified as anomalies. We integrate SHAP explainability techniques to detect inconsistencies in model decision-making, extending CSAD for SHAP-based anomaly detection. Our evaluation incorporates both anomaly detection rates with SHAP-based assessments to provide a more comprehensive measure of adversarial sample quality. We evaluate various attack strategies, examining black-box query-based and transferability-based gradient attacks across four target models. Experiments on benchmark tabular datasets reveal key differences in the attacker's risk and effort and attack quality, offering insights into the strengths, limitations, and trade-offs faced by attackers and defenders. Our findings lay the groundwork for future research on adversarial attacks and defense development in the tabular domain.

对抗攻击表格数据异常检测可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。