干净数据中隐藏的统计信号可被当作非恶意后门触发器
Statistical Adversaries: Natural Backdoor-like Adversarial Features in Clean Vision Datasets

- 从ImageNet发现与特定标签强相关的自然统计模式
- 移除随机相关性后,这些模式仍能精准操控模型预测
- 适合关注数据集漏洞与模型鲁棒性的研究者
模型特异性对抗攻击已被广泛研究。本文探讨一种不同的失效模式:视觉数据中自然存在的统计信号,虽未被恶意注入,却能表现出类似后门的触发行为。我们分析ImageNet,发现与某些标签强关联的模式;通过统计控制消除随机相关性后,验证这些信号仍能直接、可预测地改变模型输出。这些统计对手比通用噪声更具针对性,并可在不同模型架构间迁移。结果表明,部分漏洞源于数据集结构与分布,而非单一模型特性。我们得出结论:即使无数据污染,普通数据集也可能包含可被利用的对抗表面,建议在数据集审计中将虚假结构视为潜在攻击面,而不仅是偏差或可解释性问题。
原文摘要 · Abstract (English)
Model-specific adversarial attacks have been extensively studied. We study a different failure mode: naturally occurring statistical signals in vision data that can behave as backdoor-like triggers without being maliciously inserted. We call these signals statistical adversaries. We analyse ImageNet to find patterns that are strongly linked to certain labels. We then use statistical controls to remove random correlations from our candidate signals. Finally, we demonstrate that these signals directly and predictably alter model predictions. These statistical adversaries are more targeted than generic corruptions and transfer across different model architectures. This suggests that some vulnerabilities are driven by dataset structure and distribution rather than a single model's idiosyncrasies. We conclude that ordinary datasets can contain exploitable adversarial surfaces even in the absence of poisoning, and suggest that dataset audits should treat spurious structure not only as a source of bias or interpretability failure, but also as a latent attack surface for vision models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。