arXiv:2512.20815cs.CV2025-12

让摄像头和模型一起学感知,提升自动驾驶分割精度。

Learning to Sense for Driving: Joint Optics-Sensor-Model Co-Design for Semantic Segmentation

  • 从头设计光学、传感器与网络,端到端优化语义分割
  • 在KITTI-360上实现显著更高的平均交并比(mIoU)
  • 小模型低功耗运行,适合边缘部署,抗噪抗模糊

传统自动驾驶系统将摄像头设计与感知任务分离,依赖固定光学和手工设计的图像信号处理流程,优先考虑人眼可视效果而非机器语义理解。这一分离过程在去马赛克、降噪或量化阶段损失信息,迫使模型适应传感器伪影。本文提出一种面向任务的协同设计框架,将光学、传感器建模与轻量级语义分割网络统一为端到端的RAW至任务流水线。基于DeepLens[19],系统集成真实手机级镜头模型、可学习色滤阵、泊松-高斯噪声过程及量化机制,均直接针对分割目标优化。在KITTI-360上的评估显示,相比固定流水线,本方法持续提升mIoU,其中光学建模与色滤阵学习带来最大增益,尤其对细长物体或低光敏感类别表现更优。重要的是,该系统仅需约100万参数,运行速度达~28 FPS,具备边缘部署能力。视觉与定量分析进一步表明,协同设计的传感器能根据语义结构自适应采集,增强边界清晰度,并在模糊、噪声和低比特深度条件下保持准确率。这些结果确立了光学、传感器与网络全栈协同优化是构建高效、可靠、可部署感知系统的有效路径。

原文摘要 · Abstract (English)

Traditional autonomous driving pipelines decouple camera design from downstream perception, relying on fixed optics and handcrafted ISPs that prioritize human viewable imagery rather than machine semantics. This separation discards information during demosaicing, denoising, or quantization, while forcing models to adapt to sensor artifacts. We present a task-driven co-design framework that unifies optics, sensor modeling, and lightweight semantic segmentation networks into a single end-to-end RAW-to-task pipeline. Building on DeepLens[19], our system integrates realistic cellphone-scale lens models, learnable color filter arrays, Poisson-Gaussian noise processes, and quantization, all optimized directly for segmentation objectives. Evaluations on KITTI-360 show consistent mIoU improvements over fixed pipelines, with optics modeling and CFA learning providing the largest gains, especially for thin or low-light-sensitive classes. Importantly, these robustness gains are achieved with a compact ~1M-parameter model running at ~28 FPS, demonstrating edge deployability. Visual and quantitative analyses further highlight how co-designed sensors adapt acquisition to semantic structure, sharpening boundaries and maintaining accuracy under blur, noise, and low bit-depth. Together, these findings establish full-stack co-optimization of optics, sensors, and networks as a principled path toward efficient, reliable, and deployable perception in autonomous systems.

自动驾驶协同设计语义分割边缘部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。