arXiv:2507.16434cs.AIcs.LG2025-07

用元解释学习让模型驱动智能体学会无模型行为,实现自主导航。

From model-based learning to model-free behaviour with Meta-Interpretive Learning

  • 用元解释学习构建可规划的模型驱动求解器
  • 求解器训练出的控制器能解决相同导航问题
  • 适合研究自主智能体与强化学习融合的学者

“模型”是描述环境状态及智能体决策影响的理论。基于模型的智能体可预测未来行动后果并提前规划,但需完整观测环境;无模型智能体无法规划,但可在不完全观测下直接行动。能在新环境中独立运作的智能体需兼具两者能力。本文通过元解释学习,构建一个基于模型的求解器,用于训练一个可解决相同规划问题的无模型控制器。我们在两类环境中验证:随机生成迷宫和具有开阔区域的湖图。结果表明,求解器能解决的所有导航问题,控制器均可解决,二者在问题求解能力上等价。

原文摘要 · Abstract (English)

A "model" is a theory that describes the state of an environment and the effects of an agent's decisions on the environment. A model-based agent can use its model to predict the effects of its future actions and so plan ahead, but must know the state of the environment. A model-free agent cannot plan, but can act without a model and without completely observing the environment. An autonomous agent capable of acting independently in novel environments must combine both sets of capabilities. We show how to create such an agent with Meta-Interpretive Learning used to learn a model-based Solver used to train a model-free Controller that can solve the same planning problems as the Solver. We demonstrate the equivalence in problem-solving ability of the two agents on grid navigation problems in two kinds of environment: randomly generated mazes, and lake maps with wide open areas. We find that all navigation problems solved by the Solver are also solved by the Controller, indicating the two are equivalent.

元解释学习自主智能体强化学习规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。