arXiv:2411.12525cs.CVcs.AI2024-11

多视角概率融合提升分心驾驶行为定位精度

Rethinking Top Probability from Multi-view for Distracted Driver Behaviour Localization

  • 基于自监督学习模型获取多视角动作概率
  • 通过约束集成策略提升预测鲁棒性,定位更准
  • 适合关注自动驾驶安全与行为识别的研究者

自然驾驶行为定位任务旨在从真实驾驶场景的视频中识别与理解人类行为。以往研究通过识别模型结合概率后处理取得了良好表现,但模型输出的概率常含混淆信息,影响后处理效果。本文采用自监督学习框架的动作识别模型检测分心行为并生成潜在概率,利用多摄像头视角设计约束集成策略以增强预测鲁棒性,并引入条件后处理操作精确确定分心行为及时间边界。在2024年AI City Challenge赛道3的测试集A2上,本方法位列公开排行榜第六。

原文摘要 · Abstract (English)

Naturalistic driving action localization task aims to recognize and comprehend human behaviors and actions from video data captured during real-world driving scenarios. Previous studies have shown great action localization performance by applying a recognition model followed by probability-based post-processing. Nevertheless, the probabilities provided by the recognition model frequently contain confused information causing challenge for post-processing. In this work, we adopt an action recognition model based on self-supervise learning to detect distracted activities and give potential action probabilities. Subsequently, a constraint ensemble strategy takes advantages of multi-camera views to provide robust predictions. Finally, we introduce a conditional post-processing operation to locate distracted behaviours and action temporal boundaries precisely. Experimenting on test set A2, our method obtains the sixth position on the public leaderboard of track 3 of the 2024 AI City Challenge.

行为识别多视角自监督驾驶安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。