提出流式神经网络抗攻击评估新框架,提升Fuzzy ARTMAP的鲁棒性与可解释性。
Streaming Adversarial Robustness in Fuzzy ARTMAP: Mechanism-Aligned Evaluation, Progressive Training, and Interpretable Diagnostics
- 设计机制对齐的WB-Softmax攻击方法,匹配ARTMAP的类别竞争机制
- 在四个图像数据集上,原版Fuzzy ARTMAP遭攻击成功率89%-100%
- 提出渐进式训练策略,在无重放情况下实现最强流式鲁棒性
对抗鲁棒性在离线深度网络中研究充分,但对严格单遍流式神经学习器仍知之甚少。本文研究基于自适应共振理论的Fuzzy ARTMAP的对抗鲁棒性,该模型依赖类别竞争、补码编码、匹配追踪和无重放原型更新。我们提出WB-Softmax,一种与ARTMAP类别竞争和映射场预测机制对齐的可微白盒攻击代理,并形式化了流式评估原则:鲁棒性应在最终部署模型上评估。在四个图像基准上,WB-Softmax对原始Fuzzy ARTMAP的攻击成功率达89%-100%。我们发现防御策略排名会随协议变化:离线对抗训练在迁移攻击下表现良好,但在自适应白盒评估中崩溃;而渐进式两阶段选择性训练在无重放条件下提供最强整体鲁棒性。此外,ART的显式类别几何结构支持对分离坍塌和匹配得分反转的可解释诊断。这些结果为基于原型的流式学习器提供了机制对齐、协议敏感的对抗鲁棒性框架。
原文摘要 · Abstract (English)
Adversarial robustness has been studied extensively for offline deep networks, but less is known about strict single-pass streaming neural learners. This paper studies adversarial robustness in Fuzzy ARTMAP, an Adaptive Resonance Theory architecture based on category competition, complement coding, match tracking, and replay-free prototype updates. We introduce WB-Softmax, a differentiable white-box attack surrogate aligned with ARTMAP's category-competition and map-field prediction mechanism, and formalize a streaming evaluation principle requiring robustness to be assessed on the final deployed model. Across four image benchmarks, WB-Softmax achieves 89-100% attack success on vanilla Fuzzy ARTMAP models. We show that defense rankings can reverse across protocols: offline adversarial training may appear strong under transfer attacks yet collapse under adaptive white-box evaluation, whereas progressive two-stage selective training provides the strongest overall replay-free robustness. We further show that ART's explicit category geometry enables interpretable diagnosis of separation collapse and match-score inversion. These results provide a mechanism-aligned, protocol-aware framework for adversarial robustness in streaming prototype-based learners.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。