为多模态大模型设计儿童认知发展类推理测试,评估五项核心能力。
Children's Intelligence Tests Pose Challenges for MLLMs? KidGym: A 2D Grid-Based Reasoning Benchmark for MLLMs
- 基于2D网格构建12项任务,模拟儿童认知成长阶段。
- 当前主流多模态模型在规划与记忆上表现明显不足。
- 适合评估模型通用智能潜力,研究者可自定义难度和场景。
多模态大语言模型(MLLMs)融合了语言模型的语义优势与多模态数据处理能力,旨在实现更接近人类的通用智能。受韦氏儿童智力量表启发,我们提出KidGym——一个面向MLLMs的2D网格基准测试,用于评估其五大核心能力:执行、感知推理、学习、记忆与规划。该基准包含12个独特任务,每个任务至少针对一项核心能力,专门设计以衡量MLLMs的适应性与发展潜力,仿照儿童认知发展的阶段性特征。任务涵盖多样场景与随机生成的布局,确保评估的准确性和鲁棒性。KidGym支持用户完全自定义与扩展,研究人员可创建新场景并调节难度,以适应快速发展的MLLM研究需求。通过在顶尖MLLMs上应用此基准,我们揭示了模型能力的显著差异及现存局限。相关资源已公开:https://bobo-ye.github.io/KidGym/
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) combine the linguistic strengths of LLMs with the ability to process multimodal data, enbaling them to address a broader range of visual tasks. Because MLLMs aim at more general, human-like competence than language-only models, we take inspiration from the Wechsler Intelligence Scales - an established battery for evaluating children by decomposing intelligence into interpretable, testable abilities. We introduce KidGym, a comprehensive 2D grid-based benchmark for assessing five essential capabilities of MLLMs: Execution, Perception Reasoning, Learning, Memory and Planning. The benchmark comprises 12 unique tasks, each targeting at least one core capability, specifically designed to guage MLLMs' adaptability and developmental potential, mirroring the stages of children's cognitive growth. Additionally, our tasks encompass diverse scenarios and objects with randomly generated layouts, ensuring a more accurate and robust evluation of MLLM capabilities. KidGym is designed to be fully user-customizable and extensible, allowing researchers to create new evaluation scenarios and adjust difficuly levels to accommodate the rapidly growing MLLM community. Through the evaluation of state-of-the-art MLLMs using KidGym, we identified significant insights into model capabilities and revealed several limitations of current models. We release our benchmark at: https://bobo-ye.github.io/KidGym/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。