用AI生成符合国际海事规范的逼真船对船无线电对话。
Generating Realistic, Protocol-Compliant Maritime Radio Dialogues using Self-Instruct and Low-Rank Adaptation
- 通过自指导+低秩适配生成对话,嵌入26项验证确保合规性。
- 生成对话在格式、信息、逻辑上准确率超90%,多样性高。
- 适合海事AI安全系统开发,也适用于其他高危领域。
甚高频(VHF)无线电沟通失误仍是海上作业的重大安全隐患,2014至2023年间欧洲超过58%的事故由人为因素引发。尽管已长期使用,VHF通信仍受噪声、干扰、语言差异及缺乏实时转录影响,导致程序错误频发且难以纠正。构建支持实时沟通与决策的AI系统需大量高质量海事数据,但运营、法规和隐私限制使其稀缺。本研究提出一种面向合规性的自指导生成方法,可生成符合国际海事组织SMCP规范的逼真无线电对话。方法在迭代生成中集成26项验证流程,涵盖实体准确性、幻觉检测、SMCP合规性、逻辑一致性和语言多样性。采用LORA进行参数高效微调,降低训练开销,便于在资源受限的海上系统部署。为评估数据质量,引入结合自动评估与专家评审的新框架,包含格式准确性、信息准确性、唯一性和逻辑一致性四项指标。基于公开的船舶、岸基与AIS数据集实验表明,该方法生成的对话具有高度多样性、程序合规性和操作真实性。尽管下游应用如语音识别和自然语言处理暂未实现,但发布的代码、数据集与验证工具为人工智能辅助海事安全及其他高危领域提供了可复现基础。
原文摘要 · Abstract (English)
VHF radio miscommunication remains a major safety risk in maritime operations, with human factors accounting for over 58% of recorded incidents in Europe between 2014 and 2023. Despite decades of operational use, VHF radio communications are still prone to noise, interference, linguistic variability, and the absence of real-time transcription, making procedural errors both frequent and difficult to correct. Developing AI-assisted systems to support real-time communication and decision-making requires a considerable amount of high-quality maritime data, yet operational, regulatory, and privacy constraints render such datasets scarce. This study introduces a compliance aware Self-Instruct methodology for generating realistic maritime radio dialogues that conform to the IMO's SMCP. Our approach integrates a 26-filter verification pipeline directly into the iterative generation loop to enforce entity information accuracy, hallucination detection, SMCP-compliance, logical consistency, and linguistic diversity. We employ LORA for parameter-efficient fine-tuning, reducing computational overhead during training and enabling efficient deployment of the resulting models on resource-constrained maritime systems. To assess dataset quality, we introduce a novel evaluation framework combining automated and expert assessments: Format Accuracy, Information Accuracy, Uniqueness, and Logical Coherence. Experiments using publicly available vessel, coastal and AIS datasets demonstrate that the approach produces synthetically diverse, procedurally compliant, and operationally realistic dialogues. Although downstream applications such as automatic speech recognition and natural language processing are reserved for future work, the released code, datasets, and verification tools provide a reproducible foundation for artificial intelligence-assisted maritime safety and other safety-critical domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。