arXiv:2505.12656cs.CV2025-05被引 1

首个专为脉冲视频-语言对齐设计的模型,提升稀疏事件数据的理解能力。

SPKLIP: Aligning Spike Video Streams with Natural Language

  • 分层脉冲特征提取器,自适应建模多尺度时间动态
  • 脉冲-文本对比学习,实现少样本下跨模态对齐
  • 全脉冲视觉编码器节能高效,适合类脑部署

脉冲相机具备独特感知能力,但其稀疏、异步输出给语义理解带来挑战,尤其在脉冲视频-语言对齐任务中,传统模型如CLIP因模态不匹配表现不佳。我们提出SPKLIP,首个专为脉冲视频-语言对齐设计的架构。SPKLIP采用分层脉冲特征提取器,自适应建模事件流中的多尺度时间动态,并通过脉冲-文本对比学习直接对齐脉冲视频与语言,支持有效少样本学习。一种全脉冲视觉编码器变体将脉冲神经网络组件融入流水线,展现出更高的能效。实验表明,SPKLIP在基准脉冲数据集上达到领先性能,并在新贡献的真实世界数据集上展现强少样本泛化能力。其能效优势凸显了在类脑系统中部署的潜力,推动事件基多模态研究发展。源代码与数据集已公开。

原文摘要 · Abstract (English)

Spike cameras offer unique sensing capabilities but their sparse, asynchronous output challenges semantic understanding, especially for Spike Video-Language Alignment (Spike-VLA) where models like CLIP underperform due to modality mismatch. We introduce SPKLIP, the first architecture specifically for Spike-VLA. SPKLIP employs a hierarchical spike feature extractor that adaptively models multi-scale temporal dynamics in event streams, and uses spike-text contrastive learning to directly align spike video with language, enabling effective few-shot learning. A full-spiking visual encoder variant, integrating SNN components into our pipeline, demonstrates enhanced energy efficiency. Experiments show state-of-the-art performance on benchmark spike datasets and strong few-shot generalization on a newly contributed real-world dataset. SPKLIP's energy efficiency highlights its potential for neuromorphic deployment, advancing event-based multimodal research. The source code and dataset are available at [link removed for anonymity].

脉冲视觉多模态对齐类脑计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。