轻量级模型实时识别驾驶员注意力,兼顾精度与效率
SalM$^{2}$: An Extremely Lightweight Saliency Mamba Model for Real-Time Cognitive Awareness of Driver Attention
- 基于Mamba框架设计极简视觉注意模型,参数仅0.08M
- 在多个数据集上达到或超过98%的SOTA性能表现
- 适合嵌入式车载系统,适用于实时驾驶注意力监测
驾驶场景中的驾驶员注意力识别是交通场景感知技术的重要方向,旨在理解驾驶员对特定目标的关注。然而,交通场景不仅包含大量视觉信息,还蕴含与驾驶任务相关的语义信息。现有方法忽视了实际场景中的语义内容,且由于基础架构限制,模型通常参数量大、结构复杂。为此,本文提出一种基于最新Mamba框架的轻量化显著性注意力网络。如图1所示,该模型仅使用0.08M参数(仅为其他模型的0.09~11.16%),在保持SOTA性能的同时,仍能实现超过98%的基准性能,具备实时推理能力。
原文摘要 · Abstract (English)
Driver attention recognition in driving scenarios is a popular direction in traffic scene perception technology. It aims to understand human driver attention to focus on specific targets/objects in the driving scene. However, traffic scenes contain not only a large amount of visual information but also semantic information related to driving tasks. Existing methods lack attention to the actual semantic information present in driving scenes. Additionally, the traffic scene is a complex and dynamic process that requires constant attention to objects related to the current driving task. Existing models, influenced by their foundational frameworks, tend to have large parameter counts and complex structures. Therefore, this paper proposes a real-time saliency Mamba network based on the latest Mamba framework. As shown in Figure 1, our model uses very few parameters (0.08M, only 0.09~11.16% of other models), while maintaining SOTA performance or achieving over 98% of the SOTA model's performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。