让基础模型在世界模型中落地,实现无奖励的开放任务决策。
FOUNDER: Grounding Foundation Models in World Models for Open-Ended Embodied Decision Making
- 用映射函数将基础模型表征对齐到世界模型状态空间。
- 通过想象目标状态学习策略,用预测时间距离作奖励信号。
- 适合处理复杂观测或领域差异下的多任务视觉控制问题。
基础模型(FMs)与世界模型(WMs)在不同层次上具备互补的任务泛化能力。本文提出FOUNDER框架,将FMs中的可泛化知识与WMs的动态建模能力结合,实现无奖励条件下在具身环境中的开放任务求解。我们学习一个映射函数,将FM表征锚定在WM状态空间中,从而从外部观测中推断代理在世界模拟器中的物理状态。该映射使行为学习期间可通过想象学习目标条件策略,映射后的任务作为目标状态。方法利用预测到达目标的时间距离作为信息丰富的奖励信号。FOUNDER在多个多任务离线视觉控制基准上表现优异,尤其在文本或视频指定的深层语义任务中,对复杂观测或域差距场景具有强鲁棒性。实验验证了所学奖励函数与真实奖励的一致性。
原文摘要 · Abstract (English)
Foundation Models (FMs) and World Models (WMs) offer complementary strengths in task generalization at different levels. In this work, we propose FOUNDER, a framework that integrates the generalizable knowledge embedded in FMs with the dynamic modeling capabilities of WMs to enable open-ended task solving in embodied environments in a reward-free manner. We learn a mapping function that grounds FM representations in the WM state space, effectively inferring the agent's physical states in the world simulator from external observations. This mapping enables the learning of a goal-conditioned policy through imagination during behavior learning, with the mapped task serving as the goal state. Our method leverages the predicted temporal distance to the goal state as an informative reward signal. FOUNDER demonstrates superior performance on various multi-task offline visual control benchmarks, excelling in capturing the deep-level semantics of tasks specified by text or videos, particularly in scenarios involving complex observations or domain gaps where prior methods struggle. The consistency of our learned reward function with the ground-truth reward is also empirically validated. Our project website is https://sites.google.com/view/founder-rl.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。