arXiv:2607.28638cs.CLcs.LG2026-07

让智能体学会在状态中预测未来,提升决策能力。

Learning Stateful Predictive Knowledge From Experience

论文配图:Learning Stateful Predictive Knowledge From Experience
图 1 · 摘自论文原文
  • 用状态驱动的显式预测知识替代事后总结
  • 在多个任务上显著超越传统反思训练方法
  • 适合研究自主学习与复杂决策的学者

随着大型语言模型代理越来越多地从经验中学习,它们主要依赖轨迹级反思来提取洞察。从预测知识的角度看,这种做法基于情景回溯而非预测前瞻,导致结果脆弱且路径依赖。为此,我们提出状态性知识学习(SKL),将代理的关注点从轨迹级总结转向维持状态性知识:锚定于状态的显式、陈述性预测评估。我们首先通过一个示例展示状态性知识如何提供细粒度、增强泛化并支持知识自举。为进一步扩展该思想,我们引入两种算法:自蒸馏(SKL-SD)和强化学习(SKL-RL),使代理能自主从经验中提取状态基的预测知识,并学会利用它进行策略制定。在交互环境(WebShop、ScienceWorld)和复杂推理任务(ChessPuzzles)上的实验表明,赋予模型内在学习状态性预测知识的能力,显著优于当前基于反思的训练范式。

原文摘要 · Abstract (English)

As large language model (LLM) agents increasingly learn from experience, they primarily rely on trajectory-level reflection to extract insights. Viewed through the lens of predictive knowledge, we argue that this approach operates on episodic hindsight rather than predictive foresight, yielding brittle, path-dependent heuristics. To address this, we propose Stateful Knowledge Learning (SKL). SKL shifts the agent's focus from trajectory-level summarization to maintaining Stateful Knowledge: explicit, declarative predictive assessments anchored to state. We first demonstrate a motivating example showing how stateful knowledge provides granularity, enhances generalization, and enables knowledge bootstrapping. To further scale up the idea, we introduce two algorithms via self-distillation (SKL-SD) and reinforcement learning (SKL-RL), training agents to autonomously extract state-grounded predictive knowledge from experience and learn to leverage it for policy making. Experiments on interactive environments (WebShop, ScienceWorld) and a complex reasoning task (ChessPuzzles) demonstrate that equipping models with the inherent ability to learn stateful predictive knowledge significantly outpaces current reflection-based training paradigms.

智能体学习预测知识强化学习状态建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。