arXiv:2606.20752cs.CVcs.CR2026-06

用少量干净样本植入后门,让激光雷达目标检测模型悄悄误判特定物体

Mirage: a Clean-Label Backdoor against LiDAR 3D Object Detection

论文配图:Mirage: a Clean-Label Backdoor against LiDAR 3D Object Detection
图 1 · 摘自论文原文
  • 在训练数据中加入少量标签一致的污染样本,隐蔽植入触发模式
  • 仅0.5%污染率即达73%误判率,且正常检测性能几乎不变
  • 黑盒攻击,无需改标签,适合对自动驾驶感知系统发起隐蔽攻击

基于深度神经网络的激光雷达3D目标检测是安全关键型自动驾驶系统的核心感知组件。然而,近期研究揭示其易受后门攻击。现有攻击通常需白盒访问或修改标签,且集中于物体消失或边界框篡改等几何攻击。本文提出 Mirage,一种针对基于深度神经网络的激光雷达3D目标检测(LiDAR 3DOD)的黑盒、干净标签后门攻击。Mirage 将少量标签一致的污染样本注入训练集,使模型学习到触发模式与攻击者指定目标类别之间的恶意关联,同时保持正常训练语义。结果,被攻陷的模型在良性输入下表现正常,但在部署时会系统性地将触发物体误分类为目标类别。我们在多个主流的 LiDAR 3DOD 模型和基准数据集上评估 Mirage,实验表明其在仅 0.5% 污染率下达到 73% 的误判成功率,同时保持接近良性模型的检测性能。

原文摘要 · Abstract (English)

Deep neural network-based LiDAR 3D object detection serves as a critical perception component in safety-critical autonomous systems. However, recent studies have revealed its vulnerability to backdoor attacks. Existing attacks typically require white-box access or label modification and focus on geometric attacks such as object disappearance or bounding-box manipulation. In this paper, we present Mirage, a black-box and clean-label backdoor attack against deep neural network-based LiDAR 3DOD. Mirage injects a small number of label-consistent poisoning samples into the training set, causing the model to learn a malicious association between a trigger pattern and an attacker-chosen target class while preserving normal training semantics. As a result, the compromised model behaves normally on benign inputs yet systematically misclassifies triggered objects as the target class during deployment. We evaluate Mirage on multiple state-of-the-art LiDAR 3DOD models and benchmark datasets. Experimental results show that Mirage achieves a 73% misclassification success rate with a poisoning rate of only 0.5%, while maintaining detection performance close to that of benign models.

后门攻击激光雷达3D检测黑盒攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。