测试大模型能否从图像中精细观察并推理,而非仅依赖语言知识。
BLINK-Twice: You see, but do you observe? A Reasoning Benchmark on Visual Perception
- 基于视觉内容的深层推理任务,要求模型主动观察图像
- 20个主流多模态模型在该基准上表现普遍不足,平均准确率低于45%
- 引入对抗性图像对与推理链标注,适合评估视觉推理过程
近年来,多模态大语言模型(MLLMs)在推理能力方面取得快速进展,但现有推理评测仍主要评估语言推理,常将视觉输入视为可替换的上下文。为弥补这一差距,我们提出BLINK-Twice,一个以视觉为中心的推理评测基准,聚焦于具有挑战性的感知任务。本基准不依赖外部知识,要求模型仅通过视觉内容进行推理,推动从语言主导转向图像本体的推理范式。相比以往感知评测,它超越浅层感知(“看见”),强调细粒度观察与分析推理(“观察”)。BLINK-Twice包含三大核心组件:七类视觉推理挑战、自然对抗性图像对以强制依赖视觉内容,以及标注的推理链用于细粒度评估推理过程而非仅最终答案。我们评估了20个领先多模态模型,包括12个基础模型和8个增强推理模型。结果表明,当前模型在该基准上面临显著挑战;尽管链式思维等语言推理策略能提升性能,但常导致推理不稳定且冗余。我们发现重复观察图像可提升多数模型表现,而像o3这样的模型通过主动视觉交互展现出新范式的潜力。数据集已开源至https://github.com/PicoTrex/BLINK-Twice。
原文摘要 · Abstract (English)
Recently, Multimodal Large Language Models (MLLMs) have made rapid progress, particularly in enhancing their reasoning capabilities. However, existing reasoning benchmarks still primarily assess language-based reasoning, often treating visual input as replaceable context. To address this gap, we introduce BLINK-Twice, a vision-centric reasoning benchmark grounded in challenging perceptual tasks. Instead of relying on external knowledge, our tasks require models to reason from visual content alone, shifting the focus from language-based to image-grounded reasoning. Compared to prior perception benchmarks, it moves beyond shallow perception ("see") and requires fine-grained observation and analytical reasoning ("observe"). BLINK-Twice integrates three core components: seven types of visual challenges for testing visual reasoning, natural adversarial image pairs that enforce reliance on visual content, and annotated reasoning chains for fine-grained evaluation of the reasoning process rather than final answers alone. We evaluate 20 leading MLLMs, including 12 foundation models and 8 reasoning-enhanced models. BLINK-Twice poses a significant challenge to current models. While existing reasoning strategies in the language space-such as chain-of-thought or self-criticism can improve performance, they often result in unstable and redundant reasoning. We observe that repeated image observation improves performance across models, and active visual interaction, as demonstrated by models like o3, highlights the need for a new paradigm for vision reasoning. The dataset is publicly available at https://github.com/PicoTrex/BLINK-Twice
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。