arXiv:2606.14380cs.CV2026-06中稿 · the 2026 IEEE Inte…被引 1

预测未来视觉特征,提前预警交通事故

FLaRA: Predicting Future Latent Representations for Accident Anticipation

论文配图:FLaRA: Predicting Future Latent Representations for Accident Anticipation
图 1 · 摘自论文原文
  • 通过预测未来场景的隐向量来提前感知事故风险
  • 在Nexar等数据集上达到当前最优预警效果
  • 适合自动驾驶与智能交通系统开发者参考

从行车记录视频中预判交通事故是智能交通系统的关键挑战。现有方法通常直接将视觉上下文映射为碰撞概率,未显式建模驾驶场景的未来演变。本文提出FLaRA(用于事故预判的未来隐表示预测),一种新型预测架构,通过预报未来隐表示来转变这一范式。基于Video Joint-Embedding Predictive Architecture (V-JEPA2),模型以观测帧为条件,预测场景未来的隐特征,再对这些预测结果进行分类。为确保预测符合真实未来动态,引入联合训练目标,同时优化特征级重建损失与交叉熵分类损失。在Nexar数据集上的大量实验,以及在DAD、DADA-2000和DoTA等跨域基准上的验证表明,该方法在保持早期预警能力的同时,实现了最先进的性能。

原文摘要 · Abstract (English)

Anticipating traffic accidents from dashcam videos is a critical challenge in intelligent transportation systems. Existing methods typically map visual context directly to a collision probability without explicitly modeling the future evolution of the driving scene. In this paper we propose FLaRA (Predicting Future Latent Representations for Accident Anticipation), a novel predictive architecture that shifts this paradigm by forecasting future latent representations for accident anticipation. Building upon the Video Joint-Embedding Predictive Architecture (V-JEPA2), our model conditions a predictor network on observed context frames to predict the forthcoming latent features of the scene. A classifier then operates on these predicted future representations rather than only on past observations. To ensure these forecasts remain grounded in realistic future dynamics, we introduce a joint training objective that simultaneously optimizes an auxiliary feature-level reconstruction loss and a cross-entropy classification loss. Extensive evaluations on the Nexar dataset, alongside cross-domain validations on the DAD, DADA-2000, and DoTA benchmarks, demonstrate that our approach achieves state-of-the-art performance while maintaining realistic early warning capabilities.

事故预警视频预测隐表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。