防御车载联邦学习中的恶意标签攻击,提升道路状况识别安全性
Safeguarding Federated Learning-based Road Condition Classification
- 提出基于标签距离的量化指标,精准评估攻击风险
- 设计FLARE防御机制,通过输出层神经元分析降低攻击影响
- 在3个任务、6种基线中验证有效,适合自动驾驶安全研究者
联邦学习(FL)为隐私保护的自动驾驶摄像头道路状况分类(RCC)系统提供了新方案,利用车载分布式资源进行协同训练而不共享敏感图像。然而,此类框架面临新型威胁:目标标签翻转攻击(TLFA),即恶意车辆篡改训练标签,导致模型误判危险路况为良好,可能引发超速等安全事故。现有研究未充分关注此问题。本文首次揭示现有FL-RCC系统的脆弱性,提出基于标签距离的量化指标以评估安全风险,并设计名为FLARE的防御机制,通过分析输出层神经元实现抗攻击。在三个RCC任务、四种评估指标、六种基线及三种深度学习模型上的实验表明,TLFA严重损害模型性能,而FLARE能显著缓解攻击影响。
原文摘要 · Abstract (English)
Federated Learning (FL) has emerged as a promising solution for privacy-preserving autonomous driving, specifically camera-based Road Condition Classification (RCC) systems, harnessing distributed sensing, computing, and communication resources on board vehicles without sharing sensitive image data. However, the collaborative nature of FL-RCC frameworks introduces new vulnerabilities: Targeted Label Flipping Attacks (TLFAs), in which malicious clients (vehicles) deliberately alter their training data labels to compromise the learned model inference performance. Such attacks can, e.g., cause a vehicle to mis-classify slippery, dangerous road conditions as pristine and exceed recommended speed. However, TLFAs for FL-based RCC systems are largely missing. We address this challenge with a threefold contribution: 1) we disclose the vulnerability of existing FL-RCC systems to TLFAs; 2) we introduce a novel label-distance-based metric to precisely quantify the safety risks posed by TLFAs; and 3) we propose FLARE, a defensive mechanism leveraging neuron-wise analysis of the output layer to mitigate TLFA effects. Extensive experiments across three RCC tasks, four evaluation metrics, six baselines, and three deep learning models demonstrate both the severity of TLFAs on FL-RCC systems and the effectiveness of FLARE in mitigating the attack impact.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。