arXiv:2605.12703cs.CVcs.AI2026-05被引 1

构建多模态上下文学习基准,测试模型从图像中提取规则并推理的能力

MMCL-Bench: Multimodal Context Learning from Visual Rules, Procedures, and Evidence

论文配图:MMCL-Bench: Multimodal Context Learning from Visual Rules, Procedures, and Evidence
图 1 · 摘自论文原文
  • 通过图像、手册等多模态材料学习任务规则与流程
  • 最强模型仅正确解决不足1/3任务,表现远未达标
  • 适合研究多模态理解、视觉推理与复杂任务泛化的学者

我们提出MMCL-Bench,一个用于多模态上下文学习的基准:从视觉或混合模态教学内容中学习局部规则、操作流程和实证模式,并应用于新视觉实例。不同于纯文本上下文学习或标准多模态问答,该任务要求模型先从图像、截图、说明书、视频及帧序列中定位相关证据,再进行推理。MMCL-Bench包含102个任务,涵盖三类:规则系统应用、程序化任务执行、经验发现与归纳。我们采用严格评分标准评估前沿多模态模型,发现当前系统在鲁棒性上仍严重不足,最强模型在严格评测下仅完成不到三分之一任务。诊断性消融与错误分析表明,失败贯穿整个从上下文锚定、视觉证据提取、上下文推理到回答生成的全流程。因此,多模态上下文学习成为当前多模态模型的关键能力瓶颈。

原文摘要 · Abstract (English)

We introduce MMCL-Bench, a benchmark for multimodal context learning: learning task-local rules, procedures, and empirical patterns from visual or mixed-modality teaching context and applying them to new visual instances. Unlike text-only context learning or standard multimodal question answering, this setting requires models to recover and localize relevant evidence from images, screenshots, manuals, videos, and frame sequences before they can reason over the learned context. MMCL-Bench contains 102 tasks spanning three categories: rule system application, procedural task execution, and empirical discovery and induction. We evaluate frontier multimodal models with strict rubric-based scoring and find that current systems remain far from robust multimodal context learning, with even the strongest model solving fewer than one-third of tasks under strict evaluation. Diagnostic ablations and error analysis show that failures arise throughout the context-to-answer pipeline, including context anchoring, visual evidence extraction, context reasoning, and response construction. MMCL-Bench thus highlights multimodal context learning as an important unsolved capability bottleneck for current multimodal models.

多模态学习视觉推理上下文理解基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。