构建可验证的多图多步视觉推理基准,评估模型真实推理能力
EasyARC: Evaluating Vision Language Models on True Visual Reasoning
- 基于程序生成构建多图像多步骤推理任务
- 模型在复杂推理中暴露自纠错缺陷
- 适合评估强化学习与推理扩展能力
受ARC挑战启发,我们提出EasyARC,一个要求多图像、多步推理和自我修正的视觉语言基准。该基准通过程序生成,具备完全可验证性和可扩展性,适用于强化学习流程。生成器包含渐进难度层级,支持按任务类型和复杂度进行结构化评估。我们对主流视觉语言模型进行了测试并分析其失败模式。结果表明,EasyARC为评估视觉语言模型的真实推理能力和测试时扩展性能设立了新标准。数据集与评估代码已开源。
原文摘要 · Abstract (English)
Building on recent advances in language-based reasoning models, we explore multimodal reasoning that integrates vision and text. Existing multimodal benchmarks primarily test visual extraction combined with text-based reasoning, lacking true visual reasoning with more complex interactions between vision and language. Inspired by the ARC challenge, we introduce EasyARC, a vision-language benchmark requiring multi-image, multi-step reasoning, and self-correction. EasyARC is procedurally generated, fully verifiable, and scalable, making it ideal for reinforcement learning (RL) pipelines. The generators incorporate progressive difficulty levels, enabling structured evaluation across task types and complexities. We benchmark state-of-the-art vision-language models and analyze their failure modes. We argue that EasyARC sets a new standard for evaluating true reasoning and test-time scaling capabilities in vision-language models. We open-source our benchmark dataset and evaluation code.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。