多模态大模型跨模态技能组合能力差,需改进。
Multimodal LLMs Do Not Compose Skills Optimally Across Modalities
- 设计三类任务测试跨模态技能组合能力
- 所有模型在直接推理下表现均不佳,存在显著差距
- 思维链提示和微调效果有限,仍需深入研究
技能组合是指将已学习的技能结合以解决新任务的能力。随着神经网络在预训练中习得越来越复杂的技能,其能否有效组合尚不明确。本文聚焦多模态大语言模型(MLLM),研究其跨模态技能组合能力。我们设计了三个可顺序完成的任务,需组合两种模态相关技能来求解,并在两种设置下评估多个开源MLLM:i)直接提示模型解决任务;ii)采用两步级联推理,手动强制技能组合。即使在简单组合场景下,所有评估的MLLM均表现出显著的跨模态技能组合差距。为缓解此问题,我们探索两种策略:i)使用思维链提示显式指导模型进行技能组合;ii)特定微调方案以促进技能组合。尽管这些方法提升了性能,但差距依然明显,表明需进一步研究以提升MLLM的跨模态技能组合能力。
原文摘要 · Abstract (English)
Skill composition is the ability to combine previously learned skills to solve new tasks. As neural networks acquire increasingly complex skills during their pretraining, it is not clear how successfully they can compose them. In this paper, we focus on Multimodal Large Language Models (MLLM), and study their ability to compose skills across modalities. To this end, we design three evaluation tasks which can be solved sequentially composing two modality-dependent skills, and evaluate several open MLLMs under two main settings: i) prompting the model to directly solve the task, and ii) using a two-step cascaded inference approach, which manually enforces the composition of the two skills for a given task. Even with these straightforward compositions, we find that all evaluated MLLMs exhibit a significant cross-modality skill composition gap. To mitigate the aforementioned gap, we explore two alternatives: i) use chain-of-thought prompting to explicitly instruct MLLMs for skill composition and ii) a specific fine-tuning recipe to promote skill composition. Although those strategies improve model performance, they still exhibit significant skill composition gaps, suggesting that more research is needed to improve cross-modal skill composition in MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。