arXiv:2607.02575cs.CVcs.AI2026-07中稿 · ICML

让视觉语言模型根据上下文灵活调整判断标准,提升实际应用适应性。

Criterion-Conditional In-Context Learning: Evaluating Criterion-Shift Adaptation in Vision-Language Models

论文配图:Criterion-Conditional In-Context Learning: Evaluating Criterion-Shift Adaptation in Vision-Language Models
图 1 · 摘自论文原文
  • 通过上下文推断隐含判断标准并动态调整输出
  • 7B模型经简单训练后超越闭源模型且不降低泛化能力
  • 新评测基准支持多领域、多标准下的任务测试

视觉语言模型可通过上下文学习(ICL)在不更新参数的情况下执行新任务,其核心机制是利用支持集进行任务推导。标准ICL设定下,一旦任务确定,决策标准即固定不变。然而在真实场景中,许多任务具有稳定的高层意图,但其判断标准会随具体需求变化。为此,我们提出新的设置——准则条件上下文学习(CC-ICL),要求模型从上下文中推断潜在准则,并在任务语义不变的前提下调整预测结果。为评估该能力,我们设计两个互补指标:准则不变性与准则敏感性,分别衡量模型在准则变化下的鲁棒性与适应性。我们进一步构建了跨领域的CC-Bench基准,采用双层数据结构,在任务固定时仍能实现合法的真值变化。实验表明,多数模型存在刚性边界偏差,难以匹配隐含准则;而简单的多准则训练策略可显著缓解此问题,提升准则敏感性,使7B规模模型超越闭源模型,同时保持多模态通用性能。

原文摘要 · Abstract (English)

Vision-language models can perform new tasks without parameter updates through in-context learning (ICL), whose core mechanism is utilizing the support set for task induction. In the standard ICL setting, once the task is induced, its decision criterion remains fixed. However, in real-world applications, many tasks exhibit a stable high-level intent, while their decision criteria shift according to specific requirements. Thus, we introduce a new setting, denoted as Criterion-Conditional In-Context Learning (CC-ICL), where models must infer the latent criterion from context and adjust predictions accordingly under fixed task semantics. To evaluate this capability, we propose two complementary metrics, Criterion Invariance and Criterion Sensitivity, capturing the model's robustness and adaptability under criterion shifts. We further construct CC-Bench, a multi-domain benchmark that supports evaluation under the CC-ICL setting. By employing a dual-level data hierarchy, CC-Bench enables legitimate ground-truth variation conditioned on the active criterion even when the task remains fixed. Experiments on CC-Bench reveal that most models exhibit a rigid boundary bias, struggling to align their decisions with the latent criterion. We also find that even a simple multi-criterion training strategy can significantly reduce this bias, improving Criterion Sensitivity and enabling 7B-scale models to surpass proprietary models without degrading general multimodal performance.

视觉语言模型上下文学习准则适应评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。