用8×8掩码直接从压缩测量中做视觉任务,无需重建图像。
Vision without Images: End-to-End Computer Vision from Single Compressive Measurements
- 用8×8伪随机二值掩码实现物理可实现的压缩成像。
- 在极低光照下性能优于传统CMOS和压缩成像流程。
- 统一框架支持边缘检测、深度估计等多任务,模型更轻量。
快照压缩成像(SCI)具备高速、低带宽和低功耗的优势,但在低光和低信噪比条件下仍面临挑战。此外,高分辨率传感器的实际硬件限制导致大尺寸掩码难以使用,需采用更小的、硬件友好的设计。本文提出一种基于SCI的端到端计算机视觉框架,采用仅8×8大小的伪随机二值掩码,实现物理可行的部署。核心为CompDAE,一种基于STFormer架构的压缩去噪自编码器,能够直接从噪声压缩原始像素测量中执行边缘检测、深度估计等下游任务,无需图像重建。CompDAE采用受BackSlash启发的率约束训练策略,促进紧凑且可压缩的模型结构。共享编码器搭配轻量级任务特定解码器,构建统一的多任务平台。在多个数据集上的大量实验表明,CompDAE在复杂度显著降低的同时达到业界领先性能,尤其在传统CMOS与SCI流水线失效的超低光照条件下表现优异。
原文摘要 · Abstract (English)
Snapshot Compressed Imaging (SCI) offers high-speed, low-bandwidth, and energy-efficient image acquisition, but remains challenged by low-light and low signal-to-noise ratio (SNR) conditions. Moreover, practical hardware constraints in high-resolution sensors limit the use of large frame-sized masks, necessitating smaller, hardware-friendly designs. In this work, we present a novel SCI-based computer vision framework using pseudo-random binary masks of only 8$\times$8 in size for physically feasible implementations. At its core is CompDAE, a Compressive Denoising Autoencoder built on the STFormer architecture, designed to perform downstream tasks--such as edge detection and depth estimation--directly from noisy compressive raw pixel measurements without image reconstruction. CompDAE incorporates a rate-constrained training strategy inspired by BackSlash to promote compact, compressible models. A shared encoder paired with lightweight task-specific decoders enables a unified multi-task platform. Extensive experiments across multiple datasets demonstrate that CompDAE achieves state-of-the-art performance with significantly lower complexity, especially under ultra-low-light conditions where traditional CMOS and SCI pipelines fail.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。