arXiv:2503.04078cs.CV2025-03被引 1

通过时空感知与因果推理,提升复杂驾驶场景下的行为识别精度。

Spatial-Temporal Perception with Causal Inference for Naturalistic Driving Action Recognition

  • 融合时空特征与因果模块,从单路视频中捕捉细微动作变化。
  • 在两个公开数据集上达到当前最优性能,显著提升检测效率。
  • 适用于车载监控系统,尤其适合真实复杂背景下的行为分析。

自然驾驶行为识别对车辆舱内监控系统至关重要。然而,真实环境背景的复杂性给该任务带来巨大挑战,以往方法因难以观察细微行为差异且无法有效学习帧间特征,导致实际应用受限。本文提出一种新颖的时空感知(STP)架构,强调时间信息与关键物体间的空间关系,并引入因果解码器实现行为识别与时间动作定位。无需多模态输入,STP直接从RGB视频片段中提取时序与空间距离特征。随后,通过最大化所有因子分解顺序下的期望似然,联合编码双重特征。通过多尺度融合时序与空间特征,STP可感知复杂场景中的细微行为变化。此外,我们设计因果感知模块,探索帧间特征关系,显著提升检测效率与性能。我们在两个公开的驾驶员分心检测基准上验证了该方法的有效性,结果表明本框架达到当前最优水平。

原文摘要 · Abstract (English)

Naturalistic driving action recognition is essential for vehicle cabin monitoring systems. However, the complexity of real-world backgrounds presents significant challenges for this task, and previous approaches have struggled with practical implementation due to their limited ability to observe subtle behavioral differences and effectively learn inter-frame features from video. In this paper, we propose a novel Spatial-Temporal Perception (STP) architecture that emphasizes both temporal information and spatial relationships between key objects, incorporating a causal decoder to perform behavior recognition and temporal action localization. Without requiring multimodal input, STP directly extracts temporal and spatial distance features from RGB video clips. Subsequently, these dual features are jointly encoded by maximizing the expected likelihood across all possible permutations of the factorization order. By integrating temporal and spatial features at different scales, STP can perceive subtle behavioral changes in challenging scenarios. Additionally, we introduce a causal-aware module to explore relationships between video frame features, significantly enhancing detection efficiency and performance. We validate the effectiveness of our approach using two publicly available driver distraction detection benchmarks. The results demonstrate that our framework achieves state-of-the-art performance.

行为识别时空建模因果推理驾驶监控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。