arXiv:2607.08317cs.AI2026-07

测试大模型在人类轻松完成任务上的盲区,发现闭源模型仍优于开源模型。

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models

论文配图:Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models
图 1 · 摘自论文原文
  • 收集235个学生提问,设计贴近真实场景的多模态任务
  • 闭源模型在该基准上表现优于开源模型约10个百分点
  • 揭示当前模型在特定任务上普遍存在未被发现的弱点

现代AI模型在诸多基准上表现优异,但在人类看来几乎不费力的任务(如系鞋带、画五条腿的狗)上仍会失败。这表明现有基准可能低估了系统中的持续性盲点。本文提出$ exttt{blind-spots-bench}$,一个通过看似简单但对AI仍具挑战性的任务暴露盲点的基准。我们从人工智能课程学生中收集原始问题,清洗并标注结构化参考答案,构建包含235个样本的数据集,并提出适配该数据集的任务分类体系。进一步开发自动化评分流水线,评估包括开源与闭源语言、视觉-语言及图像生成模型在内的多种模型。分析显示,闭源前沿模型在该基准上显著优于开源模型,即使二者在现有基准上性能相近,差距仍达约10%。更细致分析表明,无单一模型在所有任务类型中占优,部分任务对所有模型仍具挑战。结果凸显$ exttt{blind-spots-bench}$作为诊断性压力测试的价值,可精准识别当前模型的具体缺陷。

原文摘要 · Abstract (English)

Modern AI models achieve strong performance on many established benchmarks, yet they still fail on tasks that humans find almost trivial, such as manipulating a string or drawing a dog with five legs. These examples suggest that existing benchmarks may under-measure persistent blind spots in current systems. We introduce $\texttt{blind-spots-bench}$, a benchmark designed to expose such blind spots through tasks that appear simple for humans but remain challenging for modern AI. We collect raw questions from students in an AI course, clean and annotate them with structured reference solutions, and propose a task taxonomy tailored to the resulting dataset of 235 samples. We further develop an automated grading pipeline to evaluate a wide range of models, including open-weight and closed-source language, vision-language, and image-generation models. Our analysis on $\texttt{blind-spots-bench}$ reveals that closed-source frontier models can substantially outperform open-weight models with even $\approx10\%$ gap, even when they attain comparable performance on existing benchmarks. A more fine-grained analysis shows that no single model dominates across all task types, and that some tasks remain challenging for all evaluated models. These results highlight the value of $\texttt{blind-spots-bench}$ as a diagnostic stress test for identifying concrete weaknesses in current modern models.

多模态模型评估盲点检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。