arXiv:2606.15306cs.LGcs.AI2026-06

构建可调控潜在结构的测试平台,研究大模型如何从任务序列中积累经验并持续进化。

LatentGym: A Testbed For Cross-Task Experiential Learning With Controllable Latent Structure

论文配图:LatentGym: A Testbed For Cross-Task Experiential Learning With Controllable Latent Structure
图 1 · 摘自论文原文
  • 设计带真实潜在变量的可控环境,模拟跨任务学习场景。
  • 分离探索与利用指标,量化智能体获取和运用知识的能力。
  • 适合研究个性化、交互式系统中模型适应性的研究人员。

我们设想能够持续学习的智能体系统:随着遇到一系列相关任务,它们应能推断出这些任务间共享的隐藏结构,并利用该结构改进未来决策。这种跨任务经验学习能力在个性化和交互式辅助等场景至关重要,但现有训练与评估框架缺乏共享且可控的潜在结构,无法衡量或解释智能体为何进步。我们提出LatentGym:一个由真实潜在变量控制任务结构的可控测试套件。其设计使我们能分离探索(智能体是否收集潜在信息)与利用(是否有效使用所获信息)的度量。我们在实证研究中回答三个问题:前沿模型为何无法跨任务适应;在相关任务序列上微调能否提升泛化能力及其来源;以及任务间反馈等设计如何影响训练动态与泛化表现。这些结果为研究大模型跨任务经验学习机制提供了受控基础,有助于设计在序列化、个性化和交互式场景中更可靠自适应的智能体。

原文摘要 · Abstract (English)

We envision continually learning agentic systems that become more useful over time: as they encounter sequences of related tasks, they should infer the hidden structure shared across those tasks and use it to improve future decisions. This cross-task experiential learning capability is pivotal in domains such as personalization and interactive assistance, but existing training/evaluation frameworks do not provide shared, controllable latent structures and cannot measure whether or why agents improve. We introduce LatentGym: a controllable suite in which each environment is organized around a ground-truth latent variable governing the structure across tasks. Our construction yields metrics that separate exploration (whether the agent's actions gather information about the latent) from exploitation (whether the agent uses what it has gathered). We demonstrate our suite on empirical studies addressing three questions: how and why frontier models fail to adapt across related tasks; whether post-training on related task sequences improves general cross-task adaptation, and where those gains come from; and how design choices such as inter-task feedback shape training dynamics and generalization. Together, these results establish a controlled foundation for studying how LLM agents learn from experience across tasks, and for designing agents that adapt more reliably in sequential, personalized, and interactive settings.

智能体学习跨任务适应潜在结构评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。