arXiv:2511.08246cs.AI2025-11中稿 · AAAI被引 5

提出新方法,精准定位并插入任务向量,提升多示例多模态上下文学习效果。

Where and What Matters: Sensitivity-Aware Task Vectors for Many-Shot Multimodal In-Context Learning

  • 基于激活差值的结构模式,自动识别最佳插入位置。
  • 在多个模型和任务上实现显著提升,平均性能优于现有方法。
  • 适合需要高效多示例学习的多模态应用开发者使用。

大型多模态模型(LMMs)展现出良好的上下文学习(ICL)能力,但扩展到多示例场景仍受限于上下文长度和高推理成本。现有基于任务向量的方法或忽略插入位置的重要性,或难以确定各位置的合适数值。为此,我们提出敏感性感知的任务向量插入框架(STV),解决“何处插入”与“插入什么”的问题。核心洞察是:查询-上下文对之间的激活差值具有稳定的结构模式,可作为可靠的插入线索。基于此,我们在每个敏感位置构建预聚类的激活库,并通过强化学习选择最优插入项。在Qwen-VL、Idefics-2等多模态模型及VizWiz、OK-VQA等任务上评估,STV展现良好有效性与泛化能力,持续优于已有任务向量方法。

原文摘要 · Abstract (English)

Large Multimodal Models (LMMs) have shown promising in-context learning (ICL) capabilities, but scaling to many-shot settings remains difficult due to limited context length and high inference cost. To address these challenges, task-vector-based methods have been explored by inserting compact representations of many-shot in-context demonstrations into model activations. However, existing task-vector-based methods either overlook the importance of where to insert task vectors or struggle to determine suitable values for each location. To this end, we propose a novel Sensitivity-aware Task Vector insertion framework (STV) to figure out where and what to insert. Our key insight is that activation deltas across query-context pairs exhibit consistent structural patterns, providing a reliable cue for insertion. Based on the identified sensitive-aware locations, we construct a pre-clustered activation bank for each location by clustering the activation values, and then apply reinforcement learning to choose the most suitable one to insert. We evaluate STV across a range of multimodal models (e.g., Qwen-VL, Idefics-2) and tasks (e.g., VizWiz, OK-VQA), demonstrating its effectiveness and showing consistent improvements over previous task-vector-based methods with strong generalization.

多模态上下文学习任务向量强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。