arXiv:2510.05006cs.CVcs.LG2025-10被引 1

用多隐状态表征提升驾驶行为识别的不确定性检测效率

Latent Uncertainty Representations for Video-based Driver Action and Intention Recognition

  • 通过变换层生成多个隐空间表示以估计不确定性
  • 在4个数据集上实现与顶尖方法相当的异常检测性能
  • 训练更高效,适合资源受限的车载系统部署

深度神经网络在资源受限的车载视频驾驶行为与意图识别中应用日益广泛。尽管末层概率深度学习(LL-PDL)方法能检测分布外(OOD)样本,但其性能不稳定。为此,我们提出在预训练模型基础上添加变换层,生成多个隐状态表示以估计不确定性。我们在四个视频驱动行为与意图识别数据集上,对比了提出的隐不确定性表示(LUR)和排斥训练的LUR(RLUR)与八种PDL方法,评估分类性能、校准度及基于不确定性的OOD检测能力。此外,我们为NuScenes数据集贡献了28,000帧级动作标签和1,194个视频级意图标签。结果表明,LUR和RLUR在分布内分类性能与其它LL-PDL方法相当;在基于不确定性的OOD检测中,LUR表现媲美最优方法,且训练更高效,调参更简便,无需马尔可夫链蒙特卡洛采样或复杂排斥训练流程。

原文摘要 · Abstract (English)

Deep neural networks (DNNs) are increasingly applied to safety-critical tasks in resource-constrained environments, such as video-based driver action and intention recognition. While last layer probabilistic deep learning (LL-PDL) methods can detect out-of-distribution (OOD) instances, their performance varies. As an alternative to last layer approaches, we propose extending pre-trained DNNs with transformation layers to produce multiple latent representations to estimate the uncertainty. We evaluate our latent uncertainty representation (LUR) and repulsively trained LUR (RLUR) approaches against eight PDL methods across four video-based driver action and intention recognition datasets, comparing classification performance, calibration, and uncertainty-based OOD detection. We also contribute 28,000 frame-level action labels and 1,194 video-level intention labels for the NuScenes dataset. Our results show that LUR and RLUR achieve comparable in-distribution classification performance to other LL-PDL approaches. For uncertainty-based OOD detection, LUR matches top-performing PDL methods while being more efficient to train and easier to tune than approaches that require Markov-Chain Monte Carlo sampling or repulsive training procedures.

驾驶行为识别不确定性估计视频理解轻量化模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。