arXiv:2608.29005cs.RO2026-08

提出首个针对纯摄像头端到端自动驾驶的退化容错评估基准

A Degradation-Tolerance Benchmark for Camera-Only End-to-End Driving

论文配图:A Degradation-Tolerance Benchmark for Camera-Only End-to-End Driving
图 1 · 摘自论文原文
  • 在图像加载时实时注入16类退化,覆盖5个严重等级
  • 发现模糊、JPEG压缩等退化在中等强度下会突然破坏驾驶规划
  • 区分信息丢失与质量下降,避免误判鲁棒性

纯摄像头端到端(E2E)自动驾驶模型即将部署,但其图像流常受模糊、噪声、低光、天气、帧丢失和内存故障等退化影响。当前的抗退化基准主要关注检测或鸟瞰感知,而非直接决定车辆行驶的规划输出。本文提出DriveDegrade,一个针对纯摄像头E2E驾驶中图像退化容忍度的基准。在图像加载器中实时注入16类退化,共五种严重等级,覆盖15个策略,并在nuScenes、NAVSIM上评估开环规划,同时以CARLA闭环作为锚点。结果表明:轻微退化几乎不影响规划,而造成破坏的退化类型存在明显的中等严重度阈值;脆弱性高度依赖退化类型——模糊、JPEG压缩和雨滴损伤最严重,而天气和比特错误可容忍至较高等级。进一步发现,平坦性能曲线可能源于信息删除而非质量下降,因此将退化分为‘降低质量’与‘移除信息’两类。规划器在信息被删除时必然失准,无论其对质量下降如何应对。该双轴分析能清晰量化‘自我状态捷径’效应,避免将无关性误认为鲁棒性。一个公开的视觉-语言-动作规划器在两轴上均表现平坦,且六摄像头全部遮蔽仅导致11.5%性能下降。

原文摘要 · Abstract (English)

Camera-only end-to-end (E2E) driving models are nearing deployment, where the camera stream is degraded by blur, noise, low light, weather, frame loss, and memory faults. How much a policy tolerates before its driving breaks is unclear. Corruption-robustness benchmarks target detection or bird's-eye-view perception, not the planning output that drives the car. We present DriveDegrade, a benchmark for image-degradation tolerance in camera-only E2E driving. Sixteen corruption families at five severities are injected on the fly inside the image loader, one operator reaching fifteen policies, and we evaluate open-loop planning on nuScenes and NAVSIM plus a CARLA closed-loop anchor. First, mild degradation barely affects planning, and the families that break it have a clear threshold at mid severity. Second, fragility is corruption-dependent: blur, JPEG, and raindrop damage planning most, while weather and bit error are tolerated far into the range. Third, a flat curve is ambiguous, so we separate corruptions that degrade the image from those that remove it. A planner that reads its camera must lose accuracy when information is deleted, whatever it does under quality loss. On these two axes the planners separate sharply, quantifying the ego-status shortcut without mistaking indifference for robustness. A released vision-language-action planner is flat on both axes, and blinding all six of its cameras costs it only 11.5 percent.

自动驾驶图像退化鲁棒性评估端到端驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。