让智能体与世界模型协同学习任务必需的最小表征。
Learning Task-Sufficient World Models by Synergizing Agentic Exploration and Structured Modeling

- 智能体主动探索,收集任务相关数据,形成自适应探索路径。
- 世界模型从交互数据中提炼出紧凑且任务充分的隐状态表征。
- 提升样本效率和跨技能、新任务的泛化能力,适用于机器人操控场景。
使用世界模型进行想象中的学习与规划,是训练决策智能体的有效范式。然而,现有方法常依赖高维潜在空间或通用视觉嵌入,保留了大量与控制无关的冗余信息,限制了效率与泛化能力。为此,我们研究如何让智能体学习到任务特定、最小且充分的表征。通过智能体与世界模型的闭环协同,结构化世界模型学习从信息丰富的交互数据中提炼任务充分的表示。智能体侧,通过自适应课程引导,主动探测环境以获取揭示任务相关潜在因素的轨迹;世界模型侧,对观测数据学习结构化表示,从中提炼出紧凑的任务充分隐状态。该协同机制实现了对任务充分隐状态的实证恢复,捕获所有控制相关因素。基于这些表示,所得到的策略在标准连续控制与机器人操作基准上,展现出更高的样本效率与泛化能力,涵盖跨技能、物-技组合及未见过的任务。
原文摘要 · Abstract (English)
Learning and planning in imagination using world models provides an effective paradigm for training agents for decision-making. However, existing approaches often rely on high-dimensional latent spaces or generic visual embeddings that retain many factors irrelevant to control, limiting efficiency and generalization across tasks. To this end, we study how agents can learn world models with representations that are task-specific, minimal, and sufficient for decision-making. We achieve this via a closed-loop synergy between the agent and the world model, in which structured world-model learning distills task-sufficient representations from informative interaction data. On the agent side, agents actively probe the environment to collect informative trajectories that expose task-relevant latent factors, guided by an adaptive curriculum. On the world-model side, we learn structured representations over observations to distill compact, task-sufficient latent states from the collected interaction data. This synergy enables the empirical recovery of task-sufficient latent representations that capture all control-relevant factors. Leveraging these representations, the resulting policies achieve improved sample efficiency and generalization, including generalization across skills, object-skill compositions, and previously unseen tasks on standard continuous-control and robotic-manipulation benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。