攻击者通过温度缩放破坏联邦学习模型置信度,不降低准确率却引发严重误判。
Temperature Scaling Attack Disrupting Model Confidence in Federated Learning
- 训练时注入温度缩放与学习率耦合,伪装成正常更新
- 置信度误差最高提升145%(CIFAR-100),准确率变化<2%
- 适用于医疗、自动驾驶等高风险场景的隐蔽攻击研究
预测置信度在关键系统中是核心控制信号,直接影响风险决策逻辑。现有联邦学习攻击多针对准确率或植入后门,我们首次将置信度校准列为独立攻击目标。提出温度缩放攻击(TSA),一种训练期攻击,在保持准确率不变的同时破坏校准性。通过在本地训练中引入温度缩放与学习率耦合,恶意更新维持良性优化行为,逃避基于准确率的监控和相似性检测。在非独立同分布设置下提供收敛性分析,表明该耦合保持标准收敛边界,但系统性扭曲置信度。三个基准测试中,TSA显著改变校准效果(如CIFAR-100上置信度误差增加145%),准确率变化小于2%,且对鲁棒聚合和事后校准防御仍有效。案例研究显示,置信度操纵可导致医疗漏诊率上升7.2倍或自动驾驶误报激增,即使准确率未变。结果表明,校准完整性是联邦学习中关键的攻击面。
原文摘要 · Abstract (English)
Predictive confidence serves as a foundational control signal in mission-critical systems, directly governing risk-aware logic such as escalation, abstention, and conservative fallback. While prior federated learning attacks predominantly target accuracy or implant backdoors, we identify confidence calibration as a distinct attack objective. We present the Temperature Scaling Attack (TSA), a training-time attack that degrades calibration while preserving accuracy. By injecting temperature scaling with learning rate-temperature coupling during local training, malicious updates maintain benign-like optimization behavior, evading accuracy-based monitoring and similarity-based detection. We provide a convergence analysis under non-IID settings, showing that this coupling preserves standard convergence bounds while systematically distorting confidence. Across three benchmarks, TSA substantially shifts calibration (e.g., 145% error increase on CIFAR-100) with <2 accuracy change, and remains effective under robust aggregation and post-hoc calibration defenses. Case studies further show that confidence manipulation can cause up to 7.2x increases in missed critical cases (healthcare) or false alarms (autonomous driving), even when accuracy is unchanged. Overall, our results establish calibration integrity as a critical attack surface in federated learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。