用多阶段训练打造专业图文描述模型,小模型性能超大模型。
Multi-Modal LLM based Image Captioning in ICT: Bridging the Gap Between General and Industry Domain
- 分三阶段训练:合成数据+专家标注+人机协作指令微调
- 70亿参数模型在指标上超越320亿参数的顶尖模型
- 适合需精准提取技术图像信息的行业用户
在信息与通信技术(ICT)领域,构建领域专用大语言模型或检索增强生成系统需大量高价值领域知识。这些知识不仅存在于文本中,也隐藏在图像中。传统方法可解析文档文本但无法理解图像;多模态大模型虽能识图,却缺乏领域知识。为此,本文提出一种多阶段渐进式训练策略,训练出面向ICT领域的图像描述模型(DICModel),并建立标准评估体系验证其性能。首先,利用Mermaid工具与大模型合成约7,000对图像-文本数据,用于第一阶段监督微调;其次,由ICT领域专家人工标注约2,000对数据进行第二阶段微调;最后,专家与大模型联合生成约1,500组视觉问答数据,用于指令微调。实验表明,仅70亿参数的DICModel在性能上优于其他320亿参数的先进模型。相较于70亿与320亿参数的最先进模型,本模型在BLEU指标上分别提升约56.8%和20.8%。在领域专家构建的客观问题上,准确率比Qwen2.5-VL 32B高出1%。结果表明,该工作可高效、准确地从图像中提取逻辑文本,有望推动多模态模型在ICT领域的应用发展。
原文摘要 · Abstract (English)
In the information and communications technology (ICT) industry, training a domain-specific large language model (LLM) or constructing a retrieval-augmented generation system requires a substantial amount of high-value domain knowledge. However, the knowledge is not only hidden in the textual modality but also in the image modality. Traditional methods can parse text from domain documents but dont have image captioning ability. Multi-modal LLM (MLLM) can understand images, but they do not have sufficient domain knowledge. To address the above issues, this paper proposes a multi-stage progressive training strategy to train a Domain-specific Image Captioning Model (DICModel) in ICT, and constructs a standard evaluation system to validate the performance of DICModel. Specifically, this work first synthesizes about 7K image-text pairs by combining the Mermaid tool and LLMs, which are used for the first-stage supervised-fine-tuning (SFT) of DICModel. Then, ICT-domain experts manually annotate about 2K image-text pairs for the second-stage SFT of DICModel. Finally, experts and LLMs jointly synthesize about 1.5K visual question answering data for the instruction-based SFT. Experimental results indicate that our DICModel with only 7B parameters performs better than other state-of-the-art models with 32B parameters. Compared to the SOTA models with 7B and 32B parameters, our DICModel increases the BLEU metric by approximately 56.8% and 20.8%, respectively. On the objective questions constructed by ICT domain experts, our DICModel outperforms Qwen2.5-VL 32B by 1% in terms of accuracy rate. In summary, this work can efficiently and accurately extract the logical text from images, which is expected to promote the development of multimodal models in the ICT domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。