用开源模型生成领域数据,提升多模态大模型在科研工业中的应用效果
On Domain-Adaptive Post-Training for Multimodal Large Language Models
- 通过生成-筛选流程自动生成领域专用图文指令数据
- 单阶段训练比传统两阶段更有效提升领域适应能力
- 在生物、食品、遥感等高影响力领域验证了方法有效性
将通用多模态大语言模型(MLLMs)适配至科学与工业等特定领域,对推动其实际应用具有重要意义。本文系统研究了基于后训练的MLLM领域自适应方法,涵盖数据合成、训练流程与任务评估。首先,仅使用开源模型,提出一种生成-过滤管道,基于领域图像-标题对构建多样化的视觉指令任务,生成数据在提升领域性能方面优于人工规则或强闭源模型合成的数据。其次,发现相较于通用MLLM常用的两阶段训练范式,单阶段训练在领域适配中更具优势。第三,在生物医学、食品、遥感等高影响力领域开展广泛实验,通过后训练多种MLLM并评估其在各类领域任务上的表现。最后,公开全部模型、代码与数据,以促进该方向的后续研究。
原文摘要 · Abstract (English)
Adapting general multimodal large language models (MLLMs) to specific domains, such as scientific and industrial fields, is highly significant in promoting their practical applications. This paper systematically investigates domain adaptation of MLLMs via post-training, focusing on data synthesis, training pipeline, and task evaluation. (1) Data Synthesis: Using only open-source models, we develop a generate-then-filter pipeline that curates diverse visual instruction tasks based on domain-specific image-caption pairs. The resulting data surpass the data synthesized by manual rules or strong closed-source models in enhancing domain-specific performance. (2) Training Pipeline: Unlike general MLLMs that typically adopt a two-stage training paradigm, we find that a single-stage approach is more effective for domain adaptation. (3) Task Evaluation: We conduct extensive experiments in high-impact domains such as biomedicine, food, and remote sensing, by post-training a variety of MLLMs and then evaluating MLLM performance on various domain-specific tasks. Finally, we fully open-source our models, code, and data to encourage future research in this area.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。