arXiv:2606.29548cs.LGcs.AI2026-06

用视觉语义分析驾驶行为,提升路口犹豫区决策预测精度

VISTA-DZ: Visual Semantic Trajectory Adaptation for Personalized Dilemma Zone Prediction

论文配图:VISTA-DZ: Visual Semantic Trajectory Adaptation for Personalized Dilemma Zone Prediction
图 1 · 摘自论文原文
  • 将轨迹转为图像,用视觉语言模型生成驾驶行为画像
  • 在真实数据上达90.22%准确率,模拟数据内高达93.26%
  • 适合智能交通、辅助驾驶系统开发人员参考

信号灯前的犹豫区中,驾驶员需在有限时间和距离内决定停车或通过,此决策关乎安全。准确预测停车/通行选择及决策时机对自适应信号控制、高级驾驶辅助系统及以人为中心的智能交通应用至关重要。然而,犹豫区行为高度依赖个体差异,相同轨迹可能因风险偏好、制动习惯和决策阈值不同而产生不同决策。现有个性化模型多依赖人工设计的标量特征,信息表达有限。本文提出VISTA-DZ框架,通过视觉-语言模型将历史轨迹转换为视觉表征并生成行为画像,以语义嵌入条件化双输出预测网络。模型融合双向GRU编码器、司机条件化的多头交叉注意力与特征逐通道线性调制,实现时序证据选择与特征自适应。在SDZ数据集与新收集的FDZ数据集上的实验表明,VISTA-DZ优于仅依赖轨迹或人工特征的基线模型,在域内模拟测试中达到93.26%准确率,跨20名保留驾驶员的平均准确率为90.22%。跨域结果还显示,该模型具备可行的零样本模拟到现实迁移能力,结合少量实地数据后更具现实泛化性。

原文摘要 · Abstract (English)

Driver decision making in the dilemma zone at signalized intersections is safety critical, as vehicles approaching a yellow signal must decide whether to stop or proceed within limited time and distance margins. Accurate prediction of both stop-go decisions and decision timing is important for adaptive signal control, advanced driver assistance systems, and human-centered intelligent transportation applications. However, dilemma zone behavior is strongly driver dependent. Similar approach trajectories may lead to different decisions across drivers because of differences in risk preference, braking habit, and decision threshold. Existing personalized models often rely on handcrafted scalar descriptors, which provide useful but limited summaries of individual behavior. This paper proposes VISTA-DZ, a semantic-profile-conditioned framework for personalized stop-go and decision-time prediction. Historical trajectories are converted into visual representations, interpreted by a vision-language model to generate behavioral profiles, and encoded as semantic embeddings to condition a dual-output prediction network. The final model combines a bidirectional GRU encoder, driver-conditioned multi-head cross-attention, and Feature-wise Linear Modulation for temporal evidence selection and feature adaptation. Experiments on the SDZ dataset and a newly collected FDZ dataset show that VISTA-DZ outperforms trajectory-only and handcrafted personalization baselines, achieving 93.26% in-domain simulation accuracy and 90.22% mean accuracy across 20 held-out simulation drivers. Cross-domain results further show feasible zero-shot simulation-to-real transfer and better real-world generalization when simulation data are combined with limited field data.

驾驶行为预测视觉语言模型个性化建模智能交通

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。