测试视觉语言模型能否理解图文间的因果顺序,发现表现远低于人类。
ISO-Bench: Benchmarking Multimodal Causal Reasoning in Visual-Language Models through Procedural Plans
- 通过图文配对设计,判断任务步骤的先后顺序来评估因果推理能力。
- 顶尖模型零样本F1仅0.57,链式思考最多提升至0.62,人类达0.98。
- 适合关注多模态因果理解、模型可解释性的研究者参考。
理解跨模态的因果关系是多模态模型在真实环境中运行的核心挑战。我们提出ISO-Bench,一个用于评估模型是否能推断视觉观察与程序文本之间因果依赖关系的基准。每个样本包含一个任务步骤的图像和一段来自操作计划的文本片段,目标是判断该视觉步骤发生在文本步骤之前还是之后。对十种前沿视觉-语言模型的评估显示表现不佳:最佳零样本F1仅为0.57,链式思考推理仅带来有限提升(最高0.62 F1),远低于人类水平(0.98 F1)。进一步分析揭示了提升多模态模型因果理解的具体方向。
原文摘要 · Abstract (English)
Understanding causal relationships across modalities is a core challenge for multimodal models operating in real-world environments. We introduce ISO-Bench, a benchmark for evaluating whether models can infer causal dependencies between visual observations and procedural text. Each example presents an image of a task step and a text snippet from a plan, with the goal of deciding whether the visual step occurs before or after the referenced text step. Evaluation results on ten frontier vision-language models show underwhelming performance: the best zero-shot F1 is only 0.57, and chain-of-thought reasoning yields only modest gains (up to 0.62 F1), largely behind humans (0.98 F1). Our analysis further highlights concrete directions for improving causal understanding in multimodal models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。