arXiv:2507.13425cs.CVcs.AI2025-07AAAI被引 1

用因果建模提升自动驾驶意图预测准确率

CaTFormer: Causal Temporal Transformer with Dynamic Contextual Fusion for Driving Intention Prediction

  • 引入动态上下文融合机制,精准对齐车内车外特征时序
  • 在Brain4Cars数据集上达到当前最优性能,显著提升预测精度
  • 适合自动驾驶系统开发与交通行为研究者参考

准确预测驾驶意图是提升人机共驾系统安全性和交互效率的关键,也是实现高级别自动驾驶的基础。然而,现有方法难以有效建模复杂的时空依赖关系以及人类驾驶行为的不可预测性。为此,我们提出CaTFormer,一种显式建模驾驶员行为与环境上下文之间因果关系的时序变换器。具体而言,CaTFormer引入新颖的互斥延迟融合(RDF)机制,实现内外部特征流的精确时序对齐;设计反事实残差编码(CRE)模块,系统性消除虚假相关性,揭示真实因果依赖;并构建创新的特征合成网络(FSN),自适应地将净化后的表示融合为连贯的时序表征。实验结果表明,CaTFormer在Brain4Cars数据集上达到当前最优表现,有效捕捉复杂因果时序依赖,同时提升意图预测的准确性与可解释性。

原文摘要 · Abstract (English)

Accurate prediction of driving intention is key to enhancing the safety and interactive efficiency of human-machine co-driving systems. It serves as a cornerstone for achieving high-level autonomous driving. However, current approaches remain inadequate for accurately modeling the complex spatiotemporal interdependencies and the unpredictable variability of human driving behavior. To address these challenges, we propose CaTFormer, a causal Temporal Transformer that explicitly models causal interactions between driver behavior and environmental context for robust intention prediction. Specifically, CaTFormer introduces a novel Reciprocal Delayed Fusion (RDF) mechanism for precise temporal alignment of interior and exterior feature streams, a Counterfactual Residual Encoding (CRE) module that systematically eliminates spurious correlations to reveal authentic causal dependencies, and an innovative Feature Synthesis Network (FSN) that adaptively synthesizes these purified representations into coherent temporal representations. Experimental results demonstrate that CaTFormer attains state-of-the-art performance on the Brain4Cars dataset. It effectively captures complex causal temporal dependencies and enhances both the accuracy and transparency of driving intention prediction.

自动驾驶意图预测因果建模Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。