让智能体通过头脑模拟解决从未见过的新任务,一次试错就成功。
Thinking agents for zero-shot generalization to qualitatively novel tasks
- 用组合环境设计新任务,训练时隐藏特定元素组合。
- 通过预思考与后思考表现差异选择训练任务,提升推理能力。
- 在真实环境中仅一次尝试即解决全新问题,零样本泛化能力强。
智能生物能解决其一生或进化中从未遇到过的真正新问题,关键在于‘思考’能力——即在无环境交互的情况下,通过心理模拟来规划和评估解决方案。为生成真正质变的新任务,同时保证可零样本求解(通过心理模拟),我们利用环境的组合特性:训练时隐去特定元素组合。测试任务基于该组合,确保完全新颖,但因每个元素及其两两交互已在训练中出现,仍可通过心理模拟求解。我们提出一种方法,训练具备世界模型的智能体,通过比较预思考与后思考的表现差异来选择任务。测试表明,该智能体能有效模拟多种情景,并利用模拟结果指导真实环境中的行为,在单次真实试验中成功完成新任务(零样本)。
原文摘要 · Abstract (English)
Intelligent organisms can solve truly novel problems which they have never encountered before, either in their lifetime or their evolution. An important component of this capacity is the ability to ``think'', that is, to mentally manipulate objects, concepts and behaviors in order to plan and evaluate possible solutions to novel problems, even without environment interaction. To generate problems that are truly qualitatively novel, while still solvable zero-shot (by mental simulation), we use the combinatorial nature of environments: we train the agent while withholding a specific combination of the environment's elements. The novel test task, based on this combination, is thus guaranteed to be truly novel, while still mentally simulable since the agent has been exposed to each individual element (and their pairwise interactions) during training. We propose a method to train agents endowed with world models to make use their mental simulation abilities, by selecting tasks based on the difference between the agent's pre-thinking and post-thinking performance. When tested on the novel, withheld problem, the resulting agent successfully simulated alternative scenarios and used the resulting information to guide its behavior in the actual environment, solving the novel task in a single real-environment trial (zero-shot).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。