用纯文本数据提升多模态模型零样本泛化能力,更高效。
MLAN: Language-Based Instruction Tuning Preserves and Transfers Knowledge in Multimodal Language Models
- 在指令微调中加入多样纯文本数据,构建强文本知识基础。
- 仅用一半训练令牌,性能与传统视觉主导方法相当。
- 适合追求高效、跨模态知识迁移的多模态模型研究者。
我们提出一种新的视觉指令微调策略,通过建立坚实的纯文本知识库,提升多模态大语言模型的零样本任务泛化能力。现有工作缺乏对指令微调阶段各模态重要性的充分实验,常使用大量视觉-语言数据而限制纯文本数据,并固定模态混合比例。通过在视觉指令微调阶段引入多样化纯文本数据,并控制性地改变视觉-语言数据比例,我们系统评估了模态的重要性。全面评估显示,以文本为主的指令微调方法在12个通用数据集上表现与传统的视觉主导混合方法相当,同时训练令牌用量减少至一半。研究发现,充分增加多样化的纯文本数据即可实现指令遵循能力和领域知识在模态间的有效迁移,且效率更高。
原文摘要 · Abstract (English)
We present a novel visual instruction tuning strategy to improve the zero-shot task generalization of multimodal large language models by building a firm text-only knowledge base. Existing work lacks sufficient experimentation on the importance of each modality in the instruction tuning stage, often using a majority of vision-language data while keeping text-only data limited and fixing mixtures of modalities. By incorporating diverse text-only data in the visual instruction tuning stage, we vary vision-language data in various controlled experiments to investigate the importance of modality in visual instruction tuning. Our comprehensive evaluation shows that the text-heavy instruction tuning approach is able to perform on-par with traditional vision-heavy mixtures on both modalities across 12 general datasets while using as low as half the total training tokens. We find that simply increasing sufficiently diverse text-only data enables transfer of instruction following ability and domain knowledge across modalities while being more efficient than the vision-language approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。