评测多模态大模型在图像描述任务上的表现与适应性。
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis
- 对比零样本与微调方法,测试模型跨领域适应能力。
- 多模态大模型零样本表现优异,但领域微调易丢失泛化能力。
- 适合关注大模型可适配性的研究者与应用开发者参考。
图像描述任务要求算法为视觉输入生成自然语言描述。近年来,图像描述研究与大型语言模型(LLMs)及多模态大模型(如GPT-4V、Gemini)的发展逐渐融合,后者将文本模型能力拓展至多模态场景。本文评估多模态大模型是否可替代传统图像描述网络,在多个图像描述基准上进行测试。研究涵盖模型的零样本能力及其通过提示学习、前缀微调和低秩适应等方法在不同语义领域中的适应性。结果表明,尽管多模态大模型在零样本下表现优异,但在特定领域微调时仍难以兼顾领域性能与通用性。研究讨论了这些发现对图像描述未来发展的启示,以及更灵活多模态大模型的构建方向。
原文摘要 · Abstract (English)
The task of image captioning demands an algorithm to generate natural language descriptions of visual inputs. Recent advancements have seen a convergence between image captioning research and the development of Large Language Models (LLMs) and Multimodal LLMs -- like GPT-4V and Gemini -- which extend the capabilities of text-only LLMs to multiple modalities. This paper investigates whether Multimodal LLMs can supplant traditional image captioning networks by evaluating their performance on various image description benchmarks. We explore both the zero-shot capabilities of these models and their adaptability to different semantic domains through fine-tuning methods, including prompt learning, prefix tuning, and low-rank adaptation. Our results demonstrate that while Multimodal LLMs achieve impressive zero-shot performance, fine-tuning for specific domains while maintaining their generalization capabilities intact remains challenging. We discuss the implications of these findings for future research in image captioning and the development of more adaptable Multimodal LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。