提出九维认知评估框架,让AI思维过程可被系统检验。
Multi-Dimensional Assessment for AI Cognition (MAAC): A Theoretical Framework for Process-Oriented Cognitive Evaluation of Text-Based AI Systems
- 基于认知科学理论构建九维评估体系,聚焦AI如何思考。
- 首次将认知负荷、记忆整合等机制纳入评估标准。
- 适合研究者与开发者用于诊断AI推理缺陷。
传统人工智能评估依赖任务准确率、鲁棒性等结果指标,难以揭示文本AI系统背后的认知过程。本文提出多维认知评估框架(MAAC),将评估重点从输出结果转向思维过程。该框架基于马尔的三层次假说、巴德利的工作记忆模型、斯韦勒的认知负荷理论等认知科学理论,定义了九个维度:认知负荷、工具执行、内容质量、记忆整合、复杂性处理、幻觉控制、知识迁移、处理效率和过程-结果对齐。通过五项理论分析验证其内在一致性与可实证性:维度与理论映射、覆盖矩阵、与现有评估的差距分析、诊断案例示范及先验关联性预测。MAAC为文本AI系统的认知评估提供了理论与操作基础,补充了现有以结果为导向的评测体系。
原文摘要 · Abstract (English)
Evaluating artificial intelligence systems has historically relied on outcome-based benchmarks that measure task accuracy, robustness, or fairness. While indispensable, these benchmarks provide limited diagnostic insight into the underlying cognitive processes that generate performance-leaving critical questions unanswered about how AI systems reason, integrate memory, manage complexity, or avoid generating false information. This paper introduces the Multi-Dimensional Assessment for AI Cognition (MAAC), a theoretically grounded framework for shifting evaluation from what text-based AI systems produce to how they think. MAAC defines nine cognitively motivated dimensions: Cognitive Load, Tool Execution, Content Quality, Memory Integration, Complexity Handling, Hallucination Control, Knowledge Transfer, Processing Efficiency, and Process-Outcome Alignment. Each dimension is grounded in established cognitive science theory-drawing on Marr's tri-level hypothesis, Baddeley's working memory model, Sweller's cognitive load theory, and unified theories of cognition. Five theoretical analyses provide initial support for the framework's coherence and empirical testability: dimension-to-theory mapping; a coverage matrix assessing breadth and non-redundancy; a formal gap analysis relative to current evaluation practice; a worked diagnostic illustration; and a set of a priori interdependency predictions for future empirical testing. MAAC provides a theoretical and operational framework for principled process-level cognitive assessment of text-based AI systems, complementing existing outcome-based benchmarks with cognitively grounded, multi-dimensional evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。