从人类认知角度梳理数学应用题的AI求解进展
Foundation of Intelligence: Review of Math Word Problems from Human Cognition Perspective
- 按人类认知能力分类,总结五种解题关键能力
- 对比神经网络与大模型在五大基准上的统一评测表现
- 首次系统分析近十年AI推理进展,适合推理研究者参考
数学应用题(MWP)是人工智能领域自1960年代以来的基础研究课题,旨在通过模拟人类认知智能来提升AI的推理能力。技术范式已从早期规则方法演进至深度学习模型,并快速迈向大语言模型。然而,该领域仍缺乏系统的分类框架和对当前发展趋势的深入讨论。本文从人类认知视角全面回顾相关研究,揭示近年AI模型如何逐步模拟人类认知能力。我们总结出五项核心认知能力:问题理解、逻辑组织、关联记忆、批判性思维与知识学习。围绕这些能力,回顾近十年两类主流模型——神经网络求解器与基于大语言模型的求解器,并分析其在复杂解题过程中的类人表现。此外,我们在五个主流基准上重新运行所有代表性求解器,实现统一性能对比。据我们所知,这是首个从人类推理认知角度系统分析过去十年重要研究并提供整体比较的综述。相关代码已开源于 https://github.com/Ljyustc/FoI-MWP。
原文摘要 · Abstract (English)
Math word problem (MWP) serves as a fundamental research topic in artificial intelligence (AI) dating back to 1960s. This research aims to advance the reasoning abilities of AI by mirroring the human-like cognitive intelligence. The mainstream technological paradigm has evolved from the early rule-based methods, to deep learning models, and is rapidly advancing towards large language models. However, the field still lacks a systematic taxonomy for the MWP survey along with a discussion of current development trends. Therefore, in this paper, we aim to comprehensively review related research in MWP solving through the lens of human cognition, to demonstrate how recent AI models are advancing in simulating human cognitive abilities. Specifically, we summarize 5 crucial cognitive abilities for MWP solving, including Problem Understanding, Logical Organization, Associative Memory, Critical Thinking, and Knowledge Learning. Focused on these abilities, we review two mainstream MWP models in recent 10 years: neural network solvers, and LLM based solvers, and discuss the core human-like abilities they demonstrated in their intricate problem-solving process. Moreover, we rerun all the representative MWP solvers and supplement their performance on 5 mainstream benchmarks for a unified comparison. To the best of our knowledge, this survey first comprehensively analyzes the influential MWP research of the past decade from the perspective of human reasoning cognition and provides an integrative overall comparison across existing approaches. We hope it can inspire further research in AI reasoning. Our repository is released on https://github.com/Ljyustc/FoI-MWP.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。