AI评估从识图到推理,推动智能系统真实能力的检验
The Artificial Intelligence Cognitive Examination: A Survey on the Evolution of Multimodal Evaluation from Recognition to Reasoning
- 从识别任务转向考察理解逻辑与推理过程的测评体系
- 新基准如MMBench等针对大模型设计,评估深层认知能力
- 适合关注AI真实智能水平与评测前沿的研究者
本综述梳理了多模态人工智能评估的发展历程,将其视为认知能力测试的逐步演进。随着传统基准趋于饱和,高分常掩盖模型根本缺陷,领域正经历范式转变:从仅测试‘看到什么’的识别任务,转向探究‘为何’和‘如何’理解的复杂推理评测。早期以ImageNet知识测试为基础,中期发展出针对快捷学习与组合泛化失败的GQA和视觉常识推理(VCR)等应用逻辑与理解类测试。如今面向强大多模态大语言模型(MLLMs),出现了如MMBench、SEED-Bench、MMMU等专家级集成评测,重点考察推理过程本身。最后,本文探索抽象思维、创造力与社会智能等未开垦领域。结论认为,AI评估不仅是数据集的历史,更是一个持续对抗性的改进过程,推动我们重新定义真正智能系统的标准。
原文摘要 · Abstract (English)
This survey paper chronicles the evolution of evaluation in multimodal artificial intelligence (AI), framing it as a progression of increasingly sophisticated "cognitive examinations." We argue that the field is undergoing a paradigm shift, moving from simple recognition tasks that test "what" a model sees, to complex reasoning benchmarks that probe "why" and "how" it understands. This evolution is driven by the saturation of older benchmarks, where high performance often masks fundamental weaknesses. We chart the journey from the foundational "knowledge tests" of the ImageNet era to the "applied logic and comprehension" exams such as GQA and Visual Commonsense Reasoning (VCR), which were designed specifically to diagnose systemic flaws such as shortcut learning and failures in compositional generalization. We then survey the current frontier of "expert-level integration" benchmarks (e.g., MMBench, SEED-Bench, MMMU) designed for today's powerful multimodal large language models (MLLMs), which increasingly evaluate the reasoning process itself. Finally, we explore the uncharted territories of evaluating abstract, creative, and social intelligence. We conclude that the narrative of AI evaluation is not merely a history of datasets, but a continuous, adversarial process of designing better examinations that, in turn, redefine our goals for creating truly intelligent systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。