用文本描述提升低光下脉冲相机图像重建质量
Rethinking High-speed Image Reconstruction Framework with Spike Camera
- 引入文本描述与无配对高清数据作为监督信号
- 在真实低光数据集上显著改善纹理清晰度与亮度平衡
- 适合需要鲁棒视觉理解的实时成像应用
脉冲相机作为新型类脑设备,通过连续脉冲流捕获高速场景,具有更低带宽和更高动态范围。然而,在低光条件下从脉冲输入重建高质量图像仍具挑战。传统学习方法依赖合成数据训练,但面对低光环境中的噪声脉冲时性能下降,主要因噪声建模不足及合成与真实数据之间的域差距,导致重建图像纹理模糊、噪声过多、亮度失衡。为此,我们提出SpikeCLIP框架,突破传统训练范式,利用CLIP模型对齐图文的能力,以场景文本描述和无配对高清数据作为监督信号。在真实低光数据集U-CALTECH和U-CIFAR上的实验表明,SpikeCLIP显著提升了纹理细节与亮度平衡,且重建图像与下游任务所需的视觉特征高度一致,增强了复杂环境下的鲁棒性与泛化能力。
原文摘要 · Abstract (English)
Spike cameras, as innovative neuromorphic devices, generate continuous spike streams to capture high-speed scenes with lower bandwidth and higher dynamic range than traditional RGB cameras. However, reconstructing high-quality images from the spike input under low-light conditions remains challenging. Conventional learning-based methods often rely on the synthetic dataset as the supervision for training. Still, these approaches falter when dealing with noisy spikes fired under the low-light environment, leading to further performance degradation in the real-world dataset. This phenomenon is primarily due to inadequate noise modelling and the domain gap between synthetic and real datasets, resulting in recovered images with unclear textures, excessive noise, and diminished brightness. To address these challenges, we introduce a novel spike-to-image reconstruction framework SpikeCLIP that goes beyond traditional training paradigms. Leveraging the CLIP model's powerful capability to align text and images, we incorporate the textual description of the captured scene and unpaired high-quality datasets as the supervision. Our experiments on real-world low-light datasets U-CALTECH and U-CIFAR demonstrate that SpikeCLIP significantly enhances texture details and the luminance balance of recovered images. Furthermore, the reconstructed images are well-aligned with the broader visual features needed for downstream tasks, ensuring more robust and versatile performance in challenging environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。