arXiv:2607.08724cs.LGcs.RO2026-07

用记忆宫殿思想让模型自适应推理控制动作,提升灵活性与可解释性。

Latent Memory Palace: Reasoning for Control as Autoregressive Variational Inference

论文配图:Latent Memory Palace: Reasoning for Control as Autoregressive Variational Inference
图 1 · 摘自论文原文
  • 将推理建模为自回归潜在空间中的变分推断,模仿人类渐进式思考。
  • 在仿真与真实场景中表现优异,测试时可动态分配计算资源。
  • 适用于需要灵活决策的机器人控制任务,尤其适合追求可解释性的研究者。

人类决策具有高度灵活性——有些行为即时执行,有些则需长期思虑。语言模型也展现出类似的自适应推理能力。然而,将这种能力迁移到连续控制策略中仍具挑战,因直接在语言空间中推理缺乏空间理解的精细度与动作精确性。本文提出潜空间记忆宫殿(Latent Memory Palace, LMP),通过类记忆宫殿的自回归潜在空间组织信息,实现迭代且自适应的信息检索。该方法将推理建模为带有自回归潜在分布的变分推断,并推导出一种可高效优化变分下界的潜在空间强化学习技术。由此得到的策略LMP-π在仿真与真实世界任务中均表现强劲,同时具备可解释的测试时计算资源自适应分配能力。此外,同一框架还生成了可变长度动作分词器LMP-tok,显著提升下游自回归策略性能。这些结果从变分推断视角,为控制领域的潜空间推理提供了新范式。

原文摘要 · Abstract (English)

Human decision-making is highly flexible -- some actions are taken immediately; others require longer deliberation. Language models have exhibited a similar capacity for adaptive "reasoning." However, transferring this capability to continuous control policies has been challenging, as directly reasoning in language space may lack the granularity for spatial understanding and precise motions. In this work, we show that reasoning for control policies can emerge by organizing information in an autoregressive latent space reminiscent of a memory palace, where retrieval is iterative and adaptive. Our method, Latent Memory Palace (LMP), formulates reasoning as variational inference with an autoregressive latent distribution. We derive a latent-space reinforcement learning technique to tractably optimize its variational lower bound. The resulting policy, LMP-$π$, achieves strong empirical performance in simulation and real-world domains while exhibiting interpretable, adaptive allocation of test-time compute. We further show that the same framework yields a variable-length action tokenizer, LMP-$\texttt{tok}$, which significantly improves the performance of downstream autoregressive policies. Together, these results present a new perspective on latent reasoning for control through the lens of variational inference.

控制推理变分推断自回归可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。