arXiv:2512.01882cs.LG2025-12被引 3

用脉冲神经网络实现车载多模态决策,高效实时。

New Spiking Architecture for Multi-Modal Decision-Making in Autonomous Vehicles

  • 脉冲神经元融合视觉、激光雷达与车辆方向信息
  • 在高速环境任务中实现实时决策,计算开销更低
  • 适合边缘设备部署,兼顾性能与效率

本文提出一种面向自动驾驶高层决策的端到端多模态强化学习框架。该框架通过基于交叉注意力机制的感知模块,整合摄像头图像、激光雷达点云及车辆航向信息。尽管现代多模态架构普遍采用变压器结构,但其高计算成本限制了在资源受限边缘环境中的部署。为此,我们设计了一种类变压器的脉冲时序感知架构,采用三值脉冲神经元实现高效的多模态融合。在高速环境中的多项任务评估表明,该方法在实时自动驾驶决策中兼具有效性与高效性。

原文摘要 · Abstract (English)

This work proposes an end-to-end multi-modal reinforcement learning framework for high-level decision-making in autonomous vehicles. The framework integrates heterogeneous sensory input, including camera images, LiDAR point clouds, and vehicle heading information, through a cross-attention transformer-based perception module. Although transformers have become the backbone of modern multi-modal architectures, their high computational cost limits their deployment in resource-constrained edge environments. To overcome this challenge, we propose a spiking temporal-aware transformer-like architecture that uses ternary spiking neurons for computationally efficient multi-modal fusion. Comprehensive evaluations across multiple tasks in the Highway Environment demonstrate the effectiveness and efficiency of the proposed approach for real-time autonomous decision-making.

自动驾驶脉冲神经网络多模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。