arXiv:2501.11795cs.CRcs.CV2025-01

证明了数据投毒攻击可被有效检测,给出数学保障与实验证据。

Provably effective detection of effective data poisoning attacks

  • 提出可区分投毒数据的统计检验方法——共形可分性测试。
  • 证明有效投毒必然导致可被检测,具备数学可识别性。
  • 在真实场景中验证检测效果,适合安全敏感应用研究者。

本文建立了数据集投毒攻击的数学精确定义,并证明:只要成功实施了有效的数据投毒,就必然能够被有效检测。基于一种名为共形可分性测试(Conformal Separability Test)的新统计检验方法,该工作提供了数学上的可识别性保证。同时,通过实验验证了该方法在现实世界中对投毒行为的充分检测能力,为防御数据污染提供了理论与实践双重支持。

原文摘要 · Abstract (English)

This paper establishes a mathematically precise definition of dataset poisoning attack and proves that the very act of effectively poisoning a dataset ensures that the attack can be effectively detected. On top of a mathematical guarantee that dataset poisoning is identifiable by a new statistical test that we call the Conformal Separability Test, we provide experimental evidence that we can adequately detect poisoning attempts in the real world.

数据安全投毒检测统计检验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。