用扩散模型重建低光高速下的稀疏脉冲流,提升视觉可读性。
Seeing the Unseen in Low-light Spike Streams
- 基于脉冲间隔增强纹理,聚合低光下的稀疏信息
- 引入控制网络生成高速场景,融合特征提升图像质量
- 适合处理低光高动态视觉任务的科研与工程人员
脉冲相机是一种具有高时间分辨率的类脑传感器,在高速视觉任务中表现优异。与传统相机不同,脉冲相机持续累积光子并输出异步脉冲流。由于数据模式独特,脉冲流需通过重建方法才能被人类视觉感知。然而,现有方法在低光高速场景下因噪声严重、信息稀疏而表现不佳。本文提出 Diff-SPK,一种基于扩散模型的重建方法,有效利用生成先验补充不同低光条件下的纹理信息。首先,采用增强脉冲间隔纹理(ETFI)模块从低光脉冲流中聚合稀疏信息;随后,经编码后的ETFI作为控制网络输入,用于生成高速场景。为进一步提升生成质量,我们在生成过程中引入基于ETFI的特征融合模块。
原文摘要 · Abstract (English)
Spike camera, a type of neuromorphic sensor with high-temporal resolution, shows great promise for high-speed visual tasks. Unlike traditional cameras, spike camera continuously accumulates photons and fires asynchronous spike streams. Due to unique data modality, spike streams require reconstruction methods to become perceptible to the human eye. However, lots of methods struggle to handle spike streams in low-light high-speed scenarios due to severe noise and sparse information. In this work, we propose Diff-SPK, a diffusion-based reconstruction method. Diff-SPK effectively leverages generative priors to supplement texture information under diverse low-light conditions. Specifically, it first employs an Enhanced Texture from Inter-spike Interval (ETFI) to aggregate sparse information from low-light spike streams. Then, the encoded ETFI by a suitable encoder serve as the input of ControlNet for high-speed scenes generation. To improve the quality of results, we introduce an ETFI-based feature fusion module during the generation process.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。