提出超稀疏采样策略,用更少数据实现高速视频压缩感知。
Sparse Transformer for Ultra-sparse Sampled Video Compressive Sensing
- 采用每位置仅一个子帧为1的超稀疏采样,大幅降低数据量。
- 在真实与仿真数据上均显著优于现有最优算法,重建质量更高。
- 适合芯片级部署,具有固定曝光时间与更高动态范围优势。
数字相机每像素耗电约0.1微焦耳,4K传感器以30帧/秒运行时功耗达20瓦。若未来实现千兆像素、100-1000帧/秒的摄像头,当前处理模式不可持续。为此,物理层压缩测量可将每像素功耗降低10-100倍。视频快照压缩成像(SCI)通过光学层高频调制提升有效帧率。常用采样策略为随机采样(RS),即每个掩码元素随机设为0或1。受图像插值(I2P)启发,本文提出超稀疏采样(USS),即在每个空间位置仅一个子帧为1,其余为0。构建数字微镜器件(DMD)编码系统验证其有效性。理想情况下,可将USS测量分解为子测量,使用I2P算法恢复高速帧。但因DMD与CCD不匹配,无法完全分解。为此提出BSTFormer:一种稀疏Transformer,融合局部块注意力、全局稀疏注意力和全局时序注意力,充分利用USS测量的稀疏性。在模拟与真实数据上广泛实验表明,本方法显著优于所有现有先进算法。此外,USS策略相比RS具有更高动态范围。从应用角度,因其固定曝光时间,是实现片上完整视频SCI系统的理想选择。
原文摘要 · Abstract (English)
Digital cameras consume ~0.1 microjoule per pixel to capture and encode video, resulting in a power usage of ~20W for a 4K sensor operating at 30 fps. Imagining gigapixel cameras operating at 100-1000 fps, the current processing model is unsustainable. To address this, physical layer compressive measurement has been proposed to reduce power consumption per pixel by 10-100X. Video Snapshot Compressive Imaging (SCI) introduces high frequency modulation in the optical sensor layer to increase effective frame rate. A commonly used sampling strategy of video SCI is Random Sampling (RS) where each mask element value is randomly set to be 0 or 1. Similarly, image inpainting (I2P) has demonstrated that images can be recovered from a fraction of the image pixels. Inspired by I2P, we propose Ultra-Sparse Sampling (USS) regime, where at each spatial location, only one sub-frame is set to 1 and all others are set to 0. We then build a Digital Micro-mirror Device (DMD) encoding system to verify the effectiveness of our USS strategy. Ideally, we can decompose the USS measurement into sub-measurements for which we can utilize I2P algorithms to recover high-speed frames. However, due to the mismatch between the DMD and CCD, the USS measurement cannot be perfectly decomposed. To this end, we propose BSTFormer, a sparse TransFormer that utilizes local Block attention, global Sparse attention, and global Temporal attention to exploit the sparsity of the USS measurement. Extensive results on both simulated and real-world data show that our method significantly outperforms all previous state-of-the-art algorithms. Additionally, an essential advantage of the USS strategy is its higher dynamic range than that of the RS strategy. Finally, from the application perspective, the USS strategy is a good choice to implement a complete video SCI system on chip due to its fixed exposure time.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。