arXiv:2506.09473cs.CV2025-06CVPR被引 5

用强化学习自动选最优多模态示例,提升视觉语言模型少样本能力

Provoking Multi-modal Few-Shot LVLM via Exploration-Exploitation In-Context Learning

  • 设计探索-利用强化学习框架,自适应融合多模态示例
  • 在4个VQA数据集上显著提升少样本性能,泛化能力更强
  • 适合做少样本视觉语言任务的科研与工程人员

上下文学习(ICL)是指令学习的主要趋势,通过提供清晰的任务引导和示例来增强大语言模型的任务理解与执行能力。本文研究了大型视觉语言模型(LVLM)上的ICL,探索多模态示例选择策略。现有方法面临两大挑战:一是依赖预定义或基于人工直觉的示例选择,难以覆盖多样任务需求,导致次优解;二是单独选择每个示例无法建模它们之间的交互,造成信息冗余。为此,我们提出一种新的探索-利用强化学习框架,通过探索多模态信息融合策略并自适应选择一组整体有效的示例。该框架使LVLM能通过自我探索持续优化示例选择,自主识别并生成最有效的上下文学习策略。实验结果验证了该方法在四个视觉问答(VQA)数据集上的优越性能,证明其能有效提升少样本LVLM的泛化能力。

原文摘要 · Abstract (English)

In-context learning (ICL), a predominant trend in instruction learning, aims at enhancing the performance of large language models by providing clear task guidance and examples, improving their capability in task understanding and execution. This paper investigates ICL on Large Vision-Language Models (LVLMs) and explores the policies of multi-modal demonstration selection. Existing research efforts in ICL face significant challenges: First, they rely on pre-defined demonstrations or heuristic selecting strategies based on human intuition, which are usually inadequate for covering diverse task requirements, leading to sub-optimal solutions; Second, individually selecting each demonstration fails in modeling the interactions between them, resulting in information redundancy. Unlike these prevailing efforts, we propose a new exploration-exploitation reinforcement learning framework, which explores policies to fuse multi-modal information and adaptively select adequate demonstrations as an integrated whole. The framework allows LVLMs to optimize themselves by continually refining their demonstrations through self-exploration, enabling the ability to autonomously identify and generate the most effective selection policies for in-context learning. Experimental results verify the superior performance of our approach on four Visual Question-Answering (VQA) datasets, demonstrating its effectiveness in enhancing the generalization capability of few-shot LVLMs.

视觉语言模型少样本学习强化学习多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。