用大模型生成训练数据,让小模型低成本实现高精度海事分析。
Multi-Model Synthetic Training for Mission-Critical Small Language Models
- 用GPT-4o和o3-mini生成2.1万条合成问答对,指导小模型训练。
- 微调后的Qwen2.5-7B在海事任务上达75%准确率,成本降低261倍。
- 适合缺乏标注数据的垂直领域,尤其海事安全与航运管理场景。
大型语言模型(LLMs)在多个领域表现出色,但在专业领域的应用受限于领域数据稀缺与复杂性。本文提出一种新方法,通过将大模型作为一次性教师,将32亿条自动识别系统(AIS)船舶轨迹数据转化为21,543条合成问答对,利用多模型生成(GPT-4o与o3-mini)避免过拟合,确保推理准确性。经微调的Qwen2.5-7B模型在海事任务中达到75%准确率,同时成本较直接使用大模型推理降低261倍。结果表明,经过恰当微调的小型模型可达到与昂贵大模型相当的性能。本研究推动了专用小模型合成数据生成的发展,为难以人工标注的领域提供了高度可复现的框架,可立即应用于海事安全、安保运营及船舶交通管理系统。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated remarkable capabilities across many domains, yet their application to specialized fields remains constrained by the scarcity and complexity of domain-specific training data. We present a novel approach that achieves a 261x cost reduction for maritime intelligence by using LLMs as one-time teachers rather than using them directly for inference. Our method transforms 3.2 billion Automatic Identification System (AIS) vessel tracking records into 21,543 synthetic question and answer pairs through multi-model generation (GPT-4o and o3-mini), preventing overfitting and ensuring accurate reasoning. The resulting fine-tuned Qwen2.5-7B model achieves 75% accuracy on maritime tasks, while being substantially cheaper than using a larger model for inference. We show that smaller, cheaper models -- when fine tuned properly -- can provide similar accuracy compared to larger models that are prohibitively expensive. Our work contributes to the growing field of synthetic dataset generation for specialized AI applications and presents a highly reproducible framework for domains where manual annotation is infeasible. Beyond expanding research in the growing field of specialized small language models, our approach has immediate applications in maritime safety, security operations, and vessel traffic management systems in various industries.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。