让大模型世界模型的隐状态可识别,提升推理可靠性。
Textual Belief States for World Models: Identifiable Representation Learning Under Strict Mediation

- 设计离散可解释的文本隐状态,强制预测仅依赖状态与动作
- 在TextWorld和ScienceWorld上实现57%的表征质量提升
- 适合关注模型可解释性与长期推理的AI研究者
部分可观测环境中的世界模型依赖于总结交互历史的隐状态表示,但现代基于大模型的架构中,预测性能无法反映表示质量,因历史信息绕过隐状态,导致隐状态不可识别。严格隐状态中介——要求预测仅依赖隐状态和动作——是经典原则,但在文本领域难以实现:文本隐状态为离散非可导,阻碍变分训练;且表达性强的大模型解码器易忽略瓶颈。本文阐明严格中介的必要性,证明其使表征质量可实证检验,而历史泄漏架构破坏此关系。提出文本隐状态,具备离散、可解释、变长特性,并引入树结构强化学习方法fGRPO,在训练中强制严格中介。在TextWorld和ScienceWorld上的实验表明,保持一步预测准确率的同时,表征质量最高提升57%,滚动推演性能提升98%,且随任务复杂度和时域增长而增强。
原文摘要 · Abstract (English)
World models in partially observed environments rely on latent representations that summarize interaction history, but in many modern LLM-based architectures predictive performance fails to reflect representation quality due to history bypass, rendering the latent state unidentifiable. Strict latent state mediation, requiring predictions to depend only on the latent state and action, is a classical principle that resolves this, but enforcing it in text-based settings is an open challenge: textual latent states are discrete and non-differentiable, precluding variational training, and expressive LLM decoders readily ignore the bottleneck. We show how to make strict mediation work in the text domain. We formalize why it is necessary, showing that strict mediation makes representation quality empirically testable while history-leaky architectures break this connection. We then introduce textual latent states, which are discrete, interpretable, and variable-length, and factorized GRPO (fGRPO), a tree-structured reinforcement learning method that enforces strict mediation during training. Experiments on TextWorld and ScienceWorld show preserved one-step prediction accuracy alongside up to 57\% gains in representation quality and 98\% improvements in rollout performance, increasing with task complexity and horizon.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。