提出一种可同时破坏检测与深度估计的3D对抗攻击方法
BiTAA: A Bi-Task Adversarial Attack for Object Detection and Depth Estimation via 3D Gaussian Splatting
- 基于3D高斯点云构建双任务对抗扰动,支持全图与局部区域攻击
- 实现可控近/远距离深度误判,且在多个模型上保持稳定优化
- 揭示检测与深度任务间不对称迁移风险,适合自动驾驶安全研究者
基于摄像头的感知对自动驾驶至关重要,但易受针对目标检测和单目深度估计的任务特异性对抗攻击。现有二维/三维攻击多局限于单一任务,缺乏诱导可控深度偏移的机制,也无标准化协议量化跨任务迁移,导致检测与深度之间的交互研究不足。我们提出BiTAA,一种基于3D高斯点云的双任务对抗攻击方法,通过单一扰动同时降低检测性能并引入可控的深度偏差。具体而言,设计了支持全图与补丁设置的双模型攻击框架,兼容主流检测器与深度估计算法,并可选使用期望-变换(EOT)以增强物理现实性。提出复合损失函数,将检测抑制与区域内感兴趣区域(ROIs)的符号化、幅度可控的对数深度偏置耦合,实现近/远距离误判控制,同时确保跨任务优化稳定性。还提出了统一评估协议,包含跨任务迁移指标与真实世界测试,结果表明存在持续的跨任务退化现象,且检测到深度与深度到检测的迁移呈现明显不对称性。实验凸显了多任务纯视觉感知的实际风险,推动自动驾驶场景中跨任务意识防御机制的发展。
原文摘要 · Abstract (English)
Camera-based perception is critical to autonomous driving yet remains vulnerable to task-specific adversarial manipulations in object detection and monocular depth estimation. Most existing 2D/3D attacks are developed in task silos, lack mechanisms to induce controllable depth bias, and offer no standardized protocol to quantify cross-task transfer, leaving the interaction between detection and depth underexplored. We present BiTAA, a bi-task adversarial attack built on 3D Gaussian Splatting that yields a single perturbation capable of simultaneously degrading detection and biasing monocular depth. Specifically, we introduce a dual-model attack framework that supports both full-image and patch settings and is compatible with common detectors and depth estimators, with optional expectation-over-transformation (EOT) for physical reality. In addition, we design a composite loss that couples detection suppression with a signed, magnitude-controlled log-depth bias within regions of interest (ROIs) enabling controllable near or far misperception while maintaining stable optimization across tasks. We also propose a unified evaluation protocol with cross-task transfer metrics and real-world evaluations, showing consistent cross-task degradation and a clear asymmetry between Det to Depth and from Depth to Det transfer. The results highlight practical risks for multi-task camera-only perception and motivate cross-task-aware defenses in autonomous driving scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。