arXiv:2411.03562cs.LGcs.AI2024-11被引 22

AI Agent K通过模拟人类学习过程,在数据科学竞赛中达到顶尖水平。

Kolb-Based Experiential Learning for Generalist Agents with Human-Level Kaggle Data Science Performance

  • 构建基于柯尔伯经验学习理论的自主学习框架,融合维果茨基最近发展区思想。
  • 在81个任务上实现全自动代码生成,获Kaggle竞赛9金8银12铜,超越98%参赛者。
  • 首次将人类认知学习机制融入AI,适合追求通用智能与自主进化的研究者。

人类专家能力源于交互、反思与内部模型更新的循环,这正是柯尔伯经验学习与维果茨基最近发展区的核心。当前人工智能系统,尤其是大语言模型代理,依赖静态预训练或固定流程,缺乏持续适应机制。近期研究表明,大型语言模型具备反思、修正与自我校正等早期认知特征,暗示了类人学习的基础。因此核心问题在于:能否设计出具有结构化、认知基础的学习能力的智能体?为此,我们提出一个融合柯尔伯学习循环与维果茨基最近发展区的计算框架。该架构分离外部环境交互与内部反思/抽象功能,支持认知根基的支架式学习——先在结构化环境中学习,再向开放域泛化。此方法使智能体得以掌握传统微调或简单反思方法难以应对的复杂任务。其潜力在真实世界Kaggle数据科学竞赛中得到充分验证:智能体Agent K在81项任务上实现全自动数据科学代码生成,获得Elo-MMR评分1694,超过本研究中前2%的顶级参赛者(约20万用户)的平均水平。其表现包括9枚金牌、8枚银牌、12枚铜牌,其中4金4银来自奖金赛事,是首个成功整合柯尔伯与维果茨基认知学习思想的AI系统,标志着迈向通用智能的重要一步。

原文摘要 · Abstract (English)

Human expertise emerges through iterative cycles of interaction, reflection, and internal model updating, which are central to cognitive theories such as Kolb's experiential learning and Vygotsky's zone of proximal development. In contrast, current AI systems, particularly LLM agents, rely on static pre-training or rigid workflows, lacking mechanisms for continual adaptation. Recent studies identified early cognitive traits in LLM agents (reflection, revision, and self-correction) suggesting foundational elements of human-like experiential learning. Thus the key question: Can we design LLM agents capable of structured, cognitively grounded learning similar to human processes? In response, we propose a computational framework of Kolb's learning cycle with Vygotsky's ZPD for autonomous agents. Our architecture separates extrinsic (environment interaction) and intrinsic (internal reflection/abstraction) functions, enabling cognitively grounded scaffolded learning, where the agent initially learns within structured environments, followed by open-ended generalisation. This approach empowers agents to master complex tasks ; domains that traditional fine-tuning or simple reflective methods could not tackle effectively. Its potential is powerfully demonstrated via direct comparison with humans in real-world Kaggle data science competitions. Learning fully automated data science code generation across 81 tasks, our system, Agent K, demonstrated the ability to perform the entire workflow autonomously, achieving an Elo-MMR score of 1694, beyond median score of the Kaggle Masters (the top 2% among 200,000 users) of our study. With 9 gold, 8 silver, and 12 bronze medals level performance - including 4 gold and 4 silver on prize-awarding competitions - Agent K is the 1st AI system to successfully integrate Kolb- and Vygotsky-inspired human cognitive learning, marking a major step toward generalist AI.

通用智能学习框架数据科学自主代理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。