arXiv:2509.13762cs.CV2025-09被引 2

让图像处理更懂任务,小模型也能高效提升视觉感知

Task-Aware Image Signal Processor for Advanced Visual Perception

  • 用轻量级调制算子动态调整图像统计,分层控制空间变换
  • 在多个夜间/白天数据集上提升检测与分割准确率,参数减少超70%
  • 适合移动端、车载等资源受限设备部署

近年来,计算机视觉领域越来越重视原始传感器数据(RAW),因其比传统低比特RGB图像包含更多信息。早期研究多聚焦于提升视觉质量,而近期工作则尝试利用RAW数据中的丰富信息来增强目标检测、分割等感知任务的性能。然而现有方法仍存在两大局限:大规模图像信号处理器(ISP)网络带来沉重计算开销;基于调优传统ISP流程的方法则受限于有限的表征能力。为此,我们提出任务感知图像信号处理(TA-ISP),一个紧凑的从RAW到RGB的框架,可为预训练视觉模型生成任务导向的表示。TA-ISP不采用密集卷积管道,而是预测一组轻量级、多尺度的调制算子,在全局、区域和像素层级作用,以重塑不同空间范围内的图像统计特征。这种分解式控制显著扩展了可表示的空间变化变换范围,同时严格控制内存占用、计算量和延迟。在多个包含白天与夜间条件的RAW域检测与分割基准上评估,TA-ISP持续提升下游精度,并大幅降低参数量与推理时间,非常适合部署于资源受限设备。

原文摘要 · Abstract (English)

In recent years, there has been a growing trend in computer vision towards exploiting RAW sensor data, which preserves richer information compared to conventional low-bit RGB images. Early studies mainly focused on enhancing visual quality, while more recent efforts aim to leverage the abundant information in RAW data to improve the performance of visual perception tasks such as object detection and segmentation. However, existing approaches still face two key limitations: large-scale ISP networks impose heavy computational overhead, while methods based on tuning traditional ISP pipelines are restricted by limited representational capacity.To address these issues, we propose Task-Aware Image Signal Processing (TA-ISP), a compact RAW-to-RGB framework that produces task-oriented representations for pretrained vision models. Instead of heavy dense convolutional pipelines, TA-ISP predicts a small set of lightweight, multi-scale modulation operators that act at global, regional, and pixel scales to reshape image statistics across different spatial extents. This factorized control significantly expands the range of spatially varying transforms that can be represented while keeping memory usage, computation, and latency tightly constrained. Evaluated on several RAW-domain detection and segmentation benchmarks under both daytime and nighttime conditions, TA-ISP consistently improves downstream accuracy while markedly reducing parameter count and inference time, making it well suited for deployment on resource-constrained devices.

图像处理视觉感知轻量化端侧部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。