提出近实时视频目标移除攻击,可隐蔽干扰交通感知系统
Fast Object Removal Attacks on Safety-Critical Video-based Perception Systems
- 用前后帧补丁融合重建攻击帧,实现精准目标删除
- 攻击成功率94.48%,目标检测率下降超97%且帧级相似度高
- 攻击隐蔽性强,现有检测模型难以识别,适合研究安全防御
基于视频感知系统的智能交通系统支持道路安全等关键应用,但攻击者可能篡改视频帧以破坏下游感知模块,危及行人安全。本文提出一种端到端的近实时定向目标移除攻击框架,包含四阶段:逐帧定位目标、从先前帧提取一致补丁、使用上下文感知透明度融合、重建攻击帧。在南卡罗来纳州车联网测试平台(SC-CVT)交叉路口实验表明,重建帧与原始帧具有高全局相似性,帧级峰值信噪比(PSNR)超过40 dB,结构相似性指数(SSIM)高于0.996。基于YOLO的检测器显示,目标检测率降低至最多97.59%,攻击成功率达94.48%。在不同检测器和分辨率下,单帧平均处理时间0.074至0.172秒,满足近实时要求。多款预训练篡改检测模型无法有效区分攻击与真实帧。结果表明,视频感知系统易受隐蔽目标移除攻击,导致安全应用性能下降。该研究为防御此类攻击提供了依据。
原文摘要 · Abstract (English)
By leveraging data from video-based perception systems, intelligent transportation systems (ITS) support safety-critical applications that improve road safety. However, adversaries may manipulate video frames to compromise downstream perception modules, causing failures in safety-critical functions and increasing risks to vulnerable road users. This paper presents a novel attack model and an end-to-end framework for near-real-time targeted object removal attack on a video-based safety-critical system. The end-to-end attack pipeline consists of four stages: localizing targets in each frame, retrieving coherent patches from earlier frames, blending them using context-aware alpha compositing, and reconstructing attacked frames. Experiments at an intersection on the South Carolina Connected Vehicle Testbed (SC-CVT) show that reconstructed frames have high global similarity to the originals, with frame-level Peak Signal to Noise Ratio (PSNR) above 40 dB and Structural Similarity Index Measure (SSIM) above 0.996. Using the YOLO-based detector, the attack reduces object detections by up to 97.59% and achieves a frame-level attack success rate of 94.48%. Across the evaluated detectors and frame resolutions, the mean execution time ranges from 0.074 to 0.172 seconds per frame on GPU hardware, indicating near-real-time performance in testing. The forensic evaluation using several pretrained tamper-detection models shows limited ability to distinguish reconstructed from authentic frames. The findings suggest that video-based perception is vulnerable to stealthy object removal attacks that can degrade the performance of safety-critical applications by reducing object detectability. These findings can help develop mitigation strategies against adversarial object removal attacks that threaten safety-critical applications, such as vision-based pedestrian safety systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。