arXiv:2502.07709cs.AI2025-02ICML被引 7

让大模型自主预测学习进度,高效探索海量目标。

MAGELLAN: Metacognitive predictions of learning progress guide autotelic LLM agents in large goal spaces

  • 用元认知机制在线预测自身能力与学习进展。
  • 在动态目标空间中实现样本高效的进度估计。
  • 适合需要自主规划的开放任务场景,如智能教育。

开放式学习代理需在巨大可能性空间中高效优先处理目标,聚焦于能最大化学习进度(LP)的任务。当大型语言模型(LLM)代理通过在线强化学习在高维且不断演化的目标空间中实现自驱式探索时,学习进度预测的关键挑战在于建模自身的胜任力,即元认知监控。传统方法要么需要大量采样,要么依赖脆弱的专家定义目标分组。本文提出MAGELLAN,一种元认知框架,使LLM代理能够在线学习预测自身胜任力与学习进度。通过捕捉目标间的语义关系,MAGELLAN实现了样本高效的LP估计,并可通过泛化动态适应演化的目标空间。在交互式学习环境中,MAGELLAN显著提升了学习进度预测效率与目标优先级排序,是唯一使代理能够完全掌握大规模且持续演化的目标空间的方法。结果表明,赋予LLM代理元认知能力进行学习进度预测,可有效将课程学习扩展至开放目标空间。

原文摘要 · Abstract (English)

Open-ended learning agents must efficiently prioritize goals in vast possibility spaces, focusing on those that maximize learning progress (LP). When such autotelic exploration is achieved by LLM agents trained with online RL in high-dimensional and evolving goal spaces, a key challenge for LP prediction is modeling one's own competence, a form of metacognitive monitoring. Traditional approaches either require extensive sampling or rely on brittle expert-defined goal groupings. We introduce MAGELLAN, a metacognitive framework that lets LLM agents learn to predict their competence and LP online. By capturing semantic relationships between goals, MAGELLAN enables sample-efficient LP estimation and dynamic adaptation to evolving goal spaces through generalization. In an interactive learning environment, we show that MAGELLAN improves LP prediction efficiency and goal prioritization, being the only method allowing the agent to fully master a large and evolving goal space. These results demonstrate how augmenting LLM agents with a metacognitive ability for LP predictions can effectively scale curriculum learning to open-ended goal spaces.

元认知自驱学习大模型强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。