融合AIS与CCTV数据,提升船舶轨迹预测精度。
CmIVTP: Cross-modal Interaction-based Vessel Trajectory Prediction for Maritime Intelligence

- 用跨模态交互机制融合船位与视觉环境信息。
- 在新构建的Maritime-MmD$^+$数据集上误差降低18.7%。
- 适合海上智能交通、港口监控等场景研究者参考。
海上智能交通系统对繁忙水道的航行安全与效率至关重要。然而,单一数据源限制导致船舶轨迹预测困难:小船的自动识别系统(AIS)数据常缺失或稀疏,而仅依赖闭路电视(CCTV)无法完整捕捉动态行为。为此,我们提出基于跨模态交互的船舶轨迹预测框架(CmIVTP),以建模船舶运动与环境约束间的复杂交互。具体地,设计目标感知场景编码器提取场景语义特征,有效捕获船-环境交互;引入跨模态交互变压器,融合AIS运动特征、CCTV环境特征与场景表征,利用跨模态注意力机制同时捕捉模态内语义与模态间交互,确保预测动态一致且环境可行。此外,通过聚类历史AIS轨迹构建船舶群体轨迹库,实现高效可扩展的候选轨迹生成。我们还构建了大型同步多模态数据集Maritime-MmD$^+$,融合AIS与CCTV视频数据,为多模态轨迹预测研究提供坚实支持。大量实验表明,CmIVTP在多模态驱动的船舶轨迹预测基准上表现更优。代码资源可在https://github.com/LouisYxLu/CmIVTP获取。
原文摘要 · Abstract (English)
Maritime intelligent transportation systems (MITS) are essential for ensuring navigation safety and efficiency in busy waterways. However, accurate vessel trajectory prediction remains challenging due to the limitations of single-source data. Automatic identification system (AIS) data is often sparse or unavailable for small vessels, while closed-circuit television (CCTV) data alone cannot fully capture dynamic vessel behavior. To mitigate these challenges, we propose a cross-modal interaction-based vessel trajectory prediction (named CmIVTP) framework to model the intricate interactions between vessel dynamics and environmental constraints. Specifically, we introduce a target-aware scene encoder to extract scene semantic features, effectively capturing vessel-environment interactions and enhancing trajectory prediction accuracy. In addition, we propose a cross-modal interaction transformer, which integrates AIS-derived motion features, CCTV-based environmental features, and scene representations. It leverages cross-modal attention mechanisms to simultaneously capture intra-modal semantics and inter-modal interactions, ensuring dynamically consistent and environmentally feasible predictions. Furthermore, we construct a vessel group trajectory bank by clustering historical AIS trajectories into representative motion patterns, providing an efficient and scalable approach for candidate trajectory generation. Additionally, we introduce the maritime multimodal dataset plus (named Maritime-MmD$^+$), a large-scale dataset that synchronizes AIS data and CCTV video data, providing robust support for multimodal trajectory prediction research. Extensive experiments demonstrate that CmIVTP achieves better performance on multimodal-driven vessel trajectory prediction benchmarks. The code resources for this work can be available at https://github.com/LouisYxLu/CmIVTP.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。