arXiv:2506.17632cs.CV2025-06

无需像素优化,用视觉图案即可攻击立体深度估计。

Pixel-Optimization-Free Patch Attack on Stereo Depth Estimation

  • 将攻击转化为视觉模式搜索,避开传统像素优化。
  • 在真实场景下实现高误差(D1-all > 0.4)且跨模型通用。
  • 适合研究物理攻击、自动驾驶安全与鲁棒性评估者。

立体深度估计(SDE)是自动驾驶等视觉系统感知环境的关键。已有工作表明SDE易受像素优化攻击,但此类方法仅限于数字、静态和特定视角场景,难以实用。本文提出两个贡献:首先,构建统一框架,将像素优化攻击扩展至特征提取、代价体构建、代价聚合和视差回归四个阶段;在九种SDE模型上,结合光照一致性等现实约束评估显示,现有攻击转移能力差。其次,提出首个无像素优化攻击方法PatchHunter,将补丁生成视为在视觉模式空间中搜索以破坏核心SDE假设,并采用强化学习策略高效发现有效且可迁移的模式。在KITTI数据集、高保真模拟器CARLA及实车部署中测试,PatchHunter在有效性和黑盒迁移性上均优于像素级攻击,在低光条件下仍使D1-all误差超过0.4,而像素级攻击误差接近0。

原文摘要 · Abstract (English)

Stereo Depth Estimation (SDE) is essential for scene perception in vision-based systems such as autonomous driving. Prior work shows SDE is vulnerable to pixel-optimization attacks, but these methods are limited to digital, static, and view-specific settings, making them impractical. This raises a central question: how to design deployable, adaptive, and transferable attacks under realistic constraints? We present two contributions to answer it. First, we build a unified framework that extends pixel-optimization attacks to four stereo-matching stages: feature extraction, cost-volume construction, cost aggregation, and disparity regression. Through systematic evaluation across nine SDE models with realistic constraints like photometric consistency, we show existing attacks suffer from poor transferability. Second, we propose PatchHunter, the first pixel-optimization-free attack. PatchHunter casts patch generation as a search in a structured space of visual patterns that disrupt core SDE assumptions, and uses a reinforcement learning policy to discover effective and transferable patterns efficiently. We evaluate PatchHunter on three levels: autonomous driving dataset, high-fidelity simulator, and real-world deployment. On KITTI, PatchHunter outperforms pixel-level attacks in both effectiveness and black-box transferability. Tests in CARLA and on vehicles with industrial-grade stereo cameras confirm robustness to physical variations. Even under challenging conditions such as low lighting, PatchHunter achieves a D1-all error above 0.4, while pixel-level attacks remain near 0.

立体深度对抗攻击自动驾驶强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。