arXiv:2505.17674cs.CV2025-05被引 1

让脉冲神经网络高效实现3D开放世界理解,性能超越同类模型。

SVL: Empowering Spiking Neural Networks for Efficient 3D Open-World Understanding

  • 设计多模态对齐与可重参数化融合框架,提升SNN泛化能力。
  • 零样本3D分类达85.4%准确率,多项任务优于已有SNN模型。
  • 适合追求能效比的3D视觉与多模态系统开发者使用。

脉冲神经网络(SNNs)为提取3D时空特征提供了低功耗途径,但现有SNN在性能上仍显著落后于人工神经网络(ANNs),主要因缺乏有效的预训练策略。这一局限导致其泛化能力弱、任务专用性强、多模态理解不足,尤其在多模态问答和零样本3D分类等挑战性任务中表现不佳。为此,我们提出基于脉冲的视觉-语言(SVL)预训练框架,使SNN在保持脉冲驱动效率的同时具备开放世界3D理解能力。SVL引入两个核心组件:(i) 多尺度三元对齐(MTA),实现跨3D、图像与文本模态的无标签三元组对比学习;(ii) 可重参数化视觉-语言融合(Rep-VLI),支持轻量推理,无需依赖大型文本编码器。大量实验表明,SVL在零样本3D分类中达到85.4%的准确率,超越先进ANN模型,并在下游任务中持续领先:3D分类提升6.1%、动态视觉传感器动作识别提升2.1%、3D检测提升1.1%、3D分割提升2.1%,同时保持优异能效。此外,SVL使SNN具备开放世界3D问答能力,部分场景下优于ANN。据我们所知,SVL是首个可扩展、通用且硬件友好的3D开放世界理解范式,有效弥合了SNN与ANN在复杂开放世界任务间的差距。代码已开源:https://github.com/bollossom/SVL。

原文摘要 · Abstract (English)

Spiking Neural Networks (SNNs) provide an energy-efficient way to extract 3D spatio-temporal features. However, existing SNNs still exhibit a significant performance gap compared to Artificial Neural Networks (ANNs) due to inadequate pre-training strategies. These limitations manifest as restricted generalization ability, task specificity, and a lack of multimodal understanding, particularly in challenging tasks such as multimodal question answering and zero-shot 3D classification. To overcome these challenges, we propose a Spike-based Vision-Language (SVL) pretraining framework that empowers SNNs with open-world 3D understanding while maintaining spike-driven efficiency. SVL introduces two key components: (i) Multi-scale Triple Alignment (MTA) for label-free triplet-based contrastive learning across 3D, image, and text modalities, and (ii) Re-parameterizable Vision-Language Integration (Rep-VLI) to enable lightweight inference without relying on large text encoders. Extensive experiments show that SVL achieves a top-1 accuracy of 85.4% in zero-shot 3D classification, surpassing advanced ANN models, and consistently outperforms prior SNNs on downstream tasks, including 3D classification (+6.1%), DVS action recognition (+2.1%), 3D detection (+1.1%), and 3D segmentation (+2.1%) with remarkable efficiency. Moreover, SVL enables SNNs to perform open-world 3D question answering, sometimes outperforming ANNs. To the best of our knowledge, SVL represents the first scalable, generalizable, and hardware-friendly paradigm for 3D open-world understanding, effectively bridging the gap between SNNs and ANNs in complex open-world understanding tasks. Code is available https://github.com/bollossom/SVL.

脉冲神经网络3D理解多模态能效

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。