用推理大模型提升船舶长期轨迹与目的地预测准确率
Towards Long-Horizon Vessel Trajectory and Destination Forecasting with Reasoning Large Language Models

- 基于可验证奖励的强化学习,让大模型理解航海逻辑
- 60天历史数据+30天预测,目的地准确率显著优于传统方法
- 40亿参数模型表现最佳,证明任务适配比模型大小更重要
长时序海上航行轨迹预测对航运管理、物流规划和海事风险分析至关重要,但月级预测仍研究不足。现有深度学习方法多聚焦短中期坐标外推,难以保证长时间跨度下的航线合理性和目的地正确性。本文探索结合推理能力的大语言模型进行船舶长期轨迹与目的地联合预测,提出基于可验证奖励强化学习(RLVR)的航海大模型后训练框架。构建了基于AIS数据的基准测试集,包含60天历史轨迹与30天预测周期,将轨迹转换为语义文本用于强化学习提示构造。RLVR通过强制物理合理性、早期加权轨迹监督及分层匹配与课程学习评估目的地正确性,实现对大模型的精准对齐。实验表明,经RLVR训练的模型在目的地相关指标上显著优于零样本大模型和典型深度学习基线。在评估的变体中,40亿参数模型表现最优,说明奖励兼容优化与任务适配能力比单纯使用更大模型(如80亿或140亿)更为关键。结果还显示,在微调数据有限条件下,LSTM仍是强基线;而基于Transformer的时空模型通常需要更大数据集和更丰富的结构化输入。本工作推动了语义化、验证对齐的航海预测发展,服务于实际运营决策。
原文摘要 · Abstract (English)
Long-horizon maritime trajectory prediction is important for shipping management, logistics planning, and maritime risk analysis, yet month-level forecasting remains insufficiently studied. Existing deep learning methods mainly focus on short- and mid-term coordinate extrapolation and often struggle to preserve route feasibility and destination correctness over extended horizons. This paper investigates joint long-horizon vessel trajectory and destination forecasting with reasoning-capable large language models, and develops a Maritime LLM post-training framework based on Reinforcement Learning with Verifiable Reward (RLVR). An AIS-based benchmark is constructed with 60-day historical trajectories and 30-day forecasting horizons, where trajectories are converted into semantic textual representations for RL prompt construction. RLVR aligns LLMs with maritime forecasting objectives by enforcing physical validity, providing early-weighted trajectory supervision, and evaluating destination correctness through hierarchical matching and curriculum learning. Experimental results show that RLVR-trained LLMs substantially improve over zero-shot LLMs and representative deep learning baselines, especially on destination-related metrics. Among the evaluated RLVR-trained variants, 4B LLMs achieve the best overall performance, suggesting that reward-compatible optimization and task-specific capacity matching are more important than simply using larger 8B or 14B LLMs. The results also show that LSTM remains a strong deep learning baseline under limited fine-tuning data, while Transformer-style spatio-temporal models typically require larger datasets and richer structured inputs. Overall, this work advances semantic, verifier-aligned maritime forecasting for operational decision support.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。