让智能体学会在状态中预测未来,提升决策能力。
Learning Stateful Predictive Knowledge From Experience

- 用状态驱动的显式预测知识替代事后总结
- 在多个任务上显著超越传统反思训练方法
- 适合研究自主学习与复杂决策的学者
随着大型语言模型代理越来越多地从经验中学习,它们主要依赖轨迹级反思来提取洞察。从预测知识的角度看,这种做法基于情景回溯而非预测前瞻,导致结果脆弱且路径依赖。为此,我们提出状态性知识学习(SKL),将代理的关注点从轨迹级总结转向维持状态性知识:锚定于状态的显式、陈述性预测评估。我们首先通过一个示例展示状态性知识如何提供细粒度、增强泛化并支持知识自举。为进一步扩展该思想,我们引入两种算法:自蒸馏(SKL-SD)和强化学习(SKL-RL),使代理能自主从经验中提取状态基的预测知识,并学会利用它进行策略制定。在交互环境(WebShop、ScienceWorld)和复杂推理任务(ChessPuzzles)上的实验表明,赋予模型内在学习状态性预测知识的能力,显著优于当前基于反思的训练范式。
原文摘要 · Abstract (English)
As large language model (LLM) agents increasingly learn from experience, they primarily rely on trajectory-level reflection to extract insights. Viewed through the lens of predictive knowledge, we argue that this approach operates on episodic hindsight rather than predictive foresight, yielding brittle, path-dependent heuristics. To address this, we propose Stateful Knowledge Learning (SKL). SKL shifts the agent's focus from trajectory-level summarization to maintaining Stateful Knowledge: explicit, declarative predictive assessments anchored to state. We first demonstrate a motivating example showing how stateful knowledge provides granularity, enhances generalization, and enables knowledge bootstrapping. To further scale up the idea, we introduce two algorithms via self-distillation (SKL-SD) and reinforcement learning (SKL-RL), training agents to autonomously extract state-grounded predictive knowledge from experience and learn to leverage it for policy making. Experiments on interactive environments (WebShop, ScienceWorld) and a complex reasoning task (ChessPuzzles) demonstrate that equipping models with the inherent ability to learn stateful predictive knowledge significantly outpaces current reflection-based training paradigms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。