arXiv:2511.13535cs.CV2025-11

攻击者通过微调颜色干扰模型解释,让预测不变但解释失效。

Accuracy is Not Enough: Poisoning Interpretability in Federated Learning via Color Skew

  • 在联邦学习中用颜色扰动误导模型注意力图
  • 解释精度下降35%但分类准确率仍超96%
  • 适合关注模型可解释性安全的研究者

随着机器学习模型在高安全场景中的广泛应用,视觉解释技术成为保障透明性的关键工具。本文揭示了一类新型攻击:在不损害模型准确率的前提下破坏其可解释性。具体而言,我们发现联邦学习中恶意客户端施加的微小颜色扰动,能将模型的显著性图从语义相关区域移开,同时保持预测结果不变。提出的色度扰动模块(Chromatic Perturbation Module)系统性地调整前景与背景间的色彩对比,破坏解释的一致性。这些扰动在训练轮次间累积,以隐蔽且持续的方式污染全局模型的内部特征归因。实验表明,标准训练流程无法有效检测或缓解解释退化,尤其在联邦学习环境下,细微颜色变化更难察觉。该攻击使Grad-CAM显著性图的峰值激活重叠度降低最高达35%,而所有测试数据集上的分类准确率均保持在96%以上。

原文摘要 · Abstract (English)

As machine learning models are increasingly deployed in safety-critical domains, visual explanation techniques have become essential tools for supporting transparency. In this work, we reveal a new class of attacks that compromise model interpretability without affecting accuracy. Specifically, we show that small color perturbations applied by adversarial clients in a federated learning setting can shift a model's saliency maps away from semantically meaningful regions while keeping the prediction unchanged. The proposed saliency-aware attack framework, called Chromatic Perturbation Module, systematically crafts adversarial examples by altering the color contrast between foreground and background in a way that disrupts explanation fidelity. These perturbations accumulate across training rounds, poisoning the global model's internal feature attributions in a stealthy and persistent manner. Our findings challenge a common assumption in model auditing that correct predictions imply faithful explanations and demonstrate that interpretability itself can be an attack surface. We evaluate this vulnerability across multiple datasets and show that standard training pipelines are insufficient to detect or mitigate explanation degradation, especially in the federated learning setting, where subtle color perturbations are harder to discern. Our attack reduces peak activation overlap in Grad-CAM explanations by up to 35% while preserving classification accuracy above 96% on all evaluated datasets.

联邦学习可解释性对抗攻击颜色扰动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。