arXiv:2601.17216cs.CVcs.AI2026-01中稿 · IEEE ICC 2026被引 1

用语义嵌入代替视频传输,实现低延迟高精度的协同碰撞预测。

Spatiotemporal Semantic V2X Framework for Cooperative Collision Prediction

  • 通过V-JEPA生成未来帧的时空语义嵌入,减少通信数据量
  • 相比原始视频传输降低四个数量级带宽需求,F1分数提升10%
  • 适合智能交通系统中实时安全预警场景

智能交通系统需要实时碰撞预测以保障道路安全并降低事故严重性。传统方法依赖路边单元(RSU)向车辆传输原始视频或高维感知数据,在车载通信带宽和延迟约束下不切实际。本文提出一种语义车联网框架:部署于RSU的摄像头利用视频联合嵌入预测架构(V-JEPA)生成未来帧的时空语义嵌入,并通过车联网链路传输至车辆;车辆端使用轻量级注意力探测器与分类器解码嵌入,实现碰撞预判。相比传输原始帧,该方案显著降低通信开销,同时保持预测精度。实验表明,在采用合适处理方法时,该框架使碰撞预测的F1分数提升10%,传输需求降低四个数量级,验证了语义车联网在智能交通系统中实现协同实时碰撞预测的潜力。

原文摘要 · Abstract (English)

Intelligent Transportation Systems (ITS) demand real-time collision prediction to ensure road safety and reduce accident severity. Conventional approaches rely on transmitting raw video or high-dimensional sensory data from roadside units (RSUs) to vehicles, which is impractical under vehicular communication bandwidth and latency constraints. In this work, we propose a semantic V2X framework in which RSU-mounted cameras generate spatiotemporal semantic embeddings of future frames using the Video Joint Embedding Predictive Architecture (V-JEPA). To evaluate the system, we construct a digital twin of an urban traffic environment enabling the generation of d verse traffic scenarios with both safe and collision events. These embeddings of the future frame, extracted from V-JEPA, capture task-relevant traffic dynamics and are transmitted via V2X links to vehicles, where a lightweight attentive probe and classifier decode them to predict imminent collisions. By transmitting only semantic embeddings instead of raw frames, the proposed system significantly reduces communication overhead while maintaining predictive accuracy. Experimental results demonstrate that the framework with an appropriate processing method achieves a 10% F1-score improvement for collision prediction while reducing transmission requirements by four orders of magnitude compared to raw video. This validates the potential of semantic V2X communication to enable cooperative, real-time collision prediction in ITS.

智能交通语义通信协同预测视觉嵌入

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。