arXiv:2505.18049cs.CV2025-05

模仿人眼分处理颜色与运动信息,提升视觉模型性能。

SpikeGen: Decoupled "Rods and Cones" Visual Representation Processing with Latent Generative Framework

  • 分离处理光流与颜色信号,结合生成模型增强表示
  • 在去模糊、帧重建等任务上显著优于传统方法
  • 适合需要高效处理稀疏视觉输入的实时系统

人类在动态环境中感知和学习视觉表征的过程极为复杂。从结构上看,人眼将视锥细胞和视杆细胞的功能解耦:视锥细胞主要负责颜色感知,而视杆细胞则专门检测运动,尤其是亮度变化。这两种不同模态的视觉信息在视觉皮层中整合并处理,从而增强视觉系统的鲁棒性。受此生物机制启发,现代硬件系统不仅包含对颜色敏感的RGB相机,还引入了对运动敏感的动态视觉系统,如脉冲相机(spike cameras)。基于这些进展,本研究旨在通过将分解的多模态视觉输入与现代隐空间生成框架相结合,模拟人类视觉系统。我们将其命名为SpikeGen。我们在多种脉冲-RGB任务上评估其性能,包括条件图像与视频去模糊、从脉冲流中密集帧重建以及高速场景新视角合成。大量实验表明,利用生成模型的隐空间操作能力,可有效协同增强不同视觉模态,缓解脉冲输入的空间稀疏性和RGB输入的时间稀疏性。

原文摘要 · Abstract (English)

The process through which humans perceive and learn visual representations in dynamic environments is highly complex. From a structural perspective, the human eye decouples the functions of cone and rod cells: cones are primarily responsible for color perception, while rods are specialized in detecting motion, particularly variations in light intensity. These two distinct modalities of visual information are integrated and processed within the visual cortex, thereby enhancing the robustness of the human visual system. Inspired by this biological mechanism, modern hardware systems have evolved to include not only color-sensitive RGB cameras but also motion-sensitive Dynamic Visual Systems, such as spike cameras. Building upon these advancements, this study seeks to emulate the human visual system by integrating decomposed multi-modal visual inputs with modern latent-space generative frameworks. We named it SpikeGen. We evaluate its performance across various spike-RGB tasks, including conditional image and video deblurring, dense frame reconstruction from spike streams, and high-speed scene novel-view synthesis. Supported by extensive experiments, we demonstrate that leveraging the latent space manipulation capabilities of generative models enables an effective synergistic enhancement of different visual modalities, addressing spatial sparsity in spike inputs and temporal sparsity in RGB inputs.

视觉建模脉冲神经网络生成模型多模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。