arXiv:2608.13385cs.CV2026-08中稿 · ECCV

研究什么任务下用简单向量就够了,什么情况需要更复杂的干预。

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL

论文配图:When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL
图 1 · 摘自论文原文
  • 提出选择-实现假说:查询从演示中选改变,模型计算决定如何实现。
  • 静态任务向量有效当变化在不同查询间共享度高,否则需更复杂干预。
  • 适用于多模态VQA任务,可指导方法选择而无需测试性能。

隐式多模态上下文学习将示例压缩为内部干预,形式从静态任务向量到查询相关的变换和注意力路由。尽管目标一致,这些方法在干预如何依赖查询、修改位置上差异显著,导致难以判断何时需额外复杂性。本文提出选择-实现假说:示例诱导出一组紧凑的内部变化,查询从中选择,而模型计算限制其实施方式。我们在控制的多模态任务中测试该理论,其中查询依赖性变化但任务原语与提示格式不变。通过对比正确示例与匹配反事实,测量显式M-ICL结构,并检验其是否预测干预行为。结果表明,静态任务向量的成功取决于演示引发的变化在不同查询间的共享程度;当显式M-ICL包含查询特定或分布结构时,局部加法扰动无法恢复,需更复杂干预。该规律扩展至自然视觉问答基准,支持无需测试性能的成本感知方法选择。研究提供统一的实证理论,明确任务向量何时足够,何时需更表达性的干预。

原文摘要 · Abstract (English)

Implicit multimodal in-context learning compresses demonstrations into internal interventions, ranging from static task vectors to query-conditioned transformations and attention routing. Despite their common goal, these methods differ substantially in how the intervention depends on the query and where it modifies the model, leaving unclear which additional complexity is necessary for a given task. We propose the Selection--Realization Hypothesis. It views demonstrations as inducing a compact family of internal changes from which the query selects, while the model's computation constrains how the selected change can be implemented. We evaluate this account using controlled multimodal tasks in which query dependence varies without changing the underlying task primitives or prompt format. By contrasting correct demonstrations with matched counterfactuals, we measure the structure of explicit M-ICL and test whether it predicts intervention behavior. We find that the success of a static task vector is closely tied to how much of the demonstration-induced change is shared across queries. Additional intervention complexity becomes useful when explicit M-ICL contains query-specific or distributed structure that a local additive shift cannot recover. These relationships extend to natural VQA benchmarks and support cost-aware method selection without access to test performance. Our results provide a unified empirical theory of when demonstrations can be compressed into a task vector and when a more expressive intervention is warranted.

多模态ICL任务向量模型机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。