arXiv:2605.17894cs.AI2026-05被引 2

用儿童智力测验标准评估AI的认知年龄匹配度。

Evaluating Cognitive Age Alignment in Interactive AI Agents

论文配图:Evaluating Cognitive Age Alignment in Interactive AI Agents
图 1 · 摘自论文原文
  • 基于儿童智力量表设计交互式评测基准
  • 发现当前AI在简单任务上仍无法模拟儿童认知水平
  • 适合关注AI认知能力评估的研究者参考

尽管代理型AI及其核心多模态大语言模型(MLLMs)在从日常生活到高级科研的多个领域展现出强大的语言与视觉推理能力,但人工智能与人类智能之间仍存在深刻差距。即便整合了强大工具和先进MLLMs,最先进的AI代理在一些看似简单的基础任务上仍频繁失败,而这些任务对儿童而言却轻而易举。受韦氏儿童智力量表(WISC)启发,我们提出首个基于心理测量学的交互式评测基准——ChildAgentEval,用于系统评估基于MLLM的代理在认知年龄上的对齐程度。该基准将不同代理的推理表现与特定年龄段的人类发展阶段进行对比,揭示了当前代理型AI在模拟特定年龄认知行为方面的可实现性与局限性。

原文摘要 · Abstract (English)

While agentic AI and its core multimodal large language models (MLLMs) have demonstrated remarkable promise in language and visual reasoning across domains ranging from daily life to advanced scientific research, a profound gap remains between artificial and human intelligence. Despite the integration of powerful tools and advanced MLLMs, state-of-the-art AI agents frequently fail at foundational, seemingly simple tasks that a child can resolve with ease. Inspired by the Wechsler Intelligence Scale for Children (WISC), we introduce ChildAgentEval, the first psychometrically grounded interactive benchmark for evaluating cognitive age alignment in MLLM-based agents. ChildAgentEval systematically compares the reasoning performance of various MLLM-based interactive agents against age-specific human developmental stages, exposing where current agentic AI systems can and cannot simulate age-specific cognitive behavior.

认知评估AI代理多模态模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。