arXiv:2503.09302cs.CReess.IV2025-03被引 11

提出防御数据投毒攻击的新方法,提升AI模型在恶意数据下的可靠性。

Detecting and Preventing Data Poisoning Attacks on AI Models

  • 结合异常检测与对抗训练,识别并过滤污染数据。
  • 实验显示攻击可使准确率下降27%(图像)和22%(欺诈检测)。
  • 集成学习增强鲁棒性,适合高安全要求的工业应用。

本文研究AI模型面临的数据投毒攻击问题,这是人工智能与网络安全领域日益严峻的挑战。随着技术系统在各行业的广泛应用,抵御对抗攻击的防御机制至关重要。研究旨在开发并评估新型检测与防御技术,涵盖理论框架与实际应用。通过文献综述、基于CIFAR-10和Insurance Claims数据集的实验验证,以及创新算法设计,提升模型对恶意数据干扰的抗性。研究探索了异常检测、鲁棒优化策略与集成学习等方法,以识别并缓解训练过程中的污染数据影响。实验表明,数据投毒会显著降低模型性能,使图像识别任务准确率下降高达27%,欺诈检测模型下降22%。所提防御机制,如统计异常检测与对抗训练,有效缓解攻击影响,平均恢复准确率15%-20%。结果还显示,集成学习进一步减少误报与漏报,增强模型韧性。

原文摘要 · Abstract (English)

This paper investigates the critical issue of data poisoning attacks on AI models, a growing concern in the ever-evolving landscape of artificial intelligence and cybersecurity. As advanced technology systems become increasingly prevalent across various sectors, the need for robust defence mechanisms against adversarial attacks becomes paramount. The study aims to develop and evaluate novel techniques for detecting and preventing data poisoning attacks, focusing on both theoretical frameworks and practical applications. Through a comprehensive literature review, experimental validation using the CIFAR-10 and Insurance Claims datasets, and the development of innovative algorithms, this paper seeks to enhance the resilience of AI models against malicious data manipulation. The study explores various methods, including anomaly detection, robust optimization strategies, and ensemble learning, to identify and mitigate the effects of poisoned data during model training. Experimental results indicate that data poisoning significantly degrades model performance, reducing classification accuracy by up to 27% in image recognition tasks (CIFAR-10) and 22% in fraud detection models (Insurance Claims dataset). The proposed defence mechanisms, including statistical anomaly detection and adversarial training, successfully mitigated poisoning effects, improving model robustness and restoring accuracy levels by an average of 15-20%. The findings further demonstrate that ensemble learning techniques provide an additional layer of resilience, reducing false positives and false negatives caused by adversarial data injections.

数据安全模型鲁棒性对抗攻击防御机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。