揭秘多模态上下文学习的三大关键影响因素
What Factors Affect Multi-Modal In-Context Learning? An In-Depth Exploration
- 用多模态检索器提升示范样本获取效果
- 示范内部排序比整体顺序更重要
- 提示词中加入引导说明可增强任务理解
近期多模态上下文学习(MM-ICL)进展迅速,无需额外参数调整即可在多种任务上取得优异表现。然而其有效性背后的机制仍不明确。本文系统探究了MM-ICL三个核心步骤:示范检索、示范排序与提示构建,基于6个视觉大语言模型和20种策略展开实验。结果表明:(1) 示范检索必须使用多模态检索器;(2) 示范内部顺序比跨示范顺序更关键;(3) 提示词中加入引导性说明能显著提升任务理解。本研究为未来优化MM-ICL策略提供了基础指导。
原文摘要 · Abstract (English)
Recently, rapid advancements in Multi-Modal In-Context Learning (MM-ICL) have achieved notable success, which is capable of achieving superior performance across various tasks without requiring additional parameter tuning. However, the underlying rules for the effectiveness of MM-ICL remain under-explored. To fill this gap, this work aims to investigate the research question: "What factors affect the performance of MM-ICL?'' To this end, we investigate extensive experiments on the three core steps of MM-ICL including demonstration retrieval, demonstration ordering, and prompt construction using 6 vision large language models and 20 strategies. Our findings highlight (1) the necessity of a multi-modal retriever for demonstration retrieval, (2) the importance of intra-demonstration ordering over inter-demonstration ordering, and (3) the enhancement of task comprehension through introductory instructions in prompts. We hope this study can serve as a foundational guide for optimizing MM-ICL strategies in future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。