发现视觉语言模型指令遵循能力在微调后下降,提示输出格式指导可缓解此问题。
Instruction-Following Evaluation of Large Vision-Language Models
- 构建含输出格式指令的新数据集,研究其对模型指令遵循的影响。
- 微调后模型指令遵循能力显著下降,仅12%的样本正确响应格式要求。
- 明确指定输出格式能提升准确性,适合需高可靠性的应用者参考。
随着大型语言模型(LLMs)的兴起,大量集成视觉能力的大型视觉语言模型(LVLMs)被提出。然而,经常见到的是,这些模型在使用常见训练数据集进行视觉指令微调后,其原本在语言模型中表现出的指令遵循能力明显下降,导致无法按预期执行任务指令。本研究定量验证了这一现象,并分析其根本原因。我们构建了新的训练数据集,重点标注输出格式是否明确。通过实验发现,使用包含输出格式说明的训练数据微调的LVLMs,其指令遵循能力显著优于未包含此类信息的模型。结果表明,在视觉指令微调阶段加入输出格式指导,有助于缓解指令遵循能力的退化。
原文摘要 · Abstract (English)
Following the initial flourishing of large language models (LLMs), there has been a surge in proposed large vision-language models (LVLMs) that integrate LLMs with vision capabilities. However, it has been observed that LVLMs, after tuning to visual instruction using commonly used training datasets, often fail to exhibit the instruction-following ability that was present in the LLM before integration, leading to results in which they do not follow task instructions as expected. This study quantitatively demonstrates that LVLMs' instruction-following ability declines after fine-tuning and analyzes its underlying causes. In particular, we constructed new training datasets highlighting whether the output format is specified. Then, we investigated how explicitly indicating the output format during fine-tuning affects LVLMs' instruction-following ability. Our quantitative evaluation confirmed that LVLMs' instruction-following ability declines after fine-tuning with commonly used datasets. Furthermore, we found that LVLMs trained with datasets, including instructions on output format, tend to follow instructions more accurately than models that do not. These findings suggest that including samples with instructions on output format during (visual) instruction tuning may help mitigate the decline in instruction-following abilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。