用强化学习让大模型更准预测船舶轨迹
ShipTraj-R1: Reinforcing Ship Trajectory Prediction in Large Language Models via Group Relative Policy Optimization
- 把船舶轨迹预测转为文本生成,动态提示冲突船信息
- 设计规则奖励机制,提升推理过程和预测精度
- 基于GRPO强化,适合航海安全与智能调度场景
近年来,强化微调技术显著提升了大语言模型(LLM)的推理能力,其中群体相对策略优化(GRPO)在多个领域表现优异。然而,将LLM应用于船舶轨迹预测仍处于探索阶段。本文提出ShipTraj-R1,一种基于LLM的新型框架,将船舶轨迹预测重构为文本到文本生成任务。首先,设计包含冲突船只轨迹信息的动态提示,引导模型实现自适应思维链(CoT)推理;其次,引入综合规则奖励机制,激励模型的推理格式与预测准确性;最后,通过领域特定提示与奖励驱动的GRPO机制对ShipTraj-R1进行强化,并以Qwen3作为模型骨干。在两个复杂且真实的海洋数据集上的大量实验表明,所提ShipTraj-R1相比最先进的深度学习与LLM基线方法,实现了最低误差。
原文摘要 · Abstract (English)
Recent advancements in reinforcement fine-tuning have significantly improved the reasoning ability of large language models (LLMs). In particular, methods such as group relative policy optimization (GRPO) have demonstrated strong capabilities across various fields. However, applying LLMs to ship trajectory prediction remains largely unexplored. In this paper, we propose ShipTraj-R1, a novel LLM-based framework that reformulates ship trajectory prediction as a text-to-text generation problem. (1) We design a dynamic prompt containing trajectory information about conflicting ships to guide the model to achieve adaptive chain-of-thought (CoT) reasoning. (2) We introduce a comprehensive rule-based reward mechanism to incentivize the reasoning format and prediction accuracy of the model. (3) Our ShipTraj-R1 is reinforced through the GRPO mechanism guided by domain-specific prompts and rewards, and utilizes the Qwen3 as the model backbone. Extensive experimental results on two complex and real-world maritime datasets show that the proposed ShipTraj-R1 achieves the least error compared with state-of-the-art deep learning and LLM-based baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。