比较了世界模型与模仿学习的控制能力,发现预测未来行为不等于更好控制。
On the Capability Separation Between World-Model Policy Learning and Imitated World-Action Models
- 将世界动作模型分解为预测未来和推断动作,对比其与直接模仿学习的差异
- 在理想条件下,两者都能复现观察行为策略,说明预测未来不提升控制能力上限
- 关键区别在于能否预测指定动作的后果,适合研究策略优化与因果推理的学者
世界动作模型通过预测未来结果再反推对应动作。尽管这种分解可提升表征学习与数据效率,但其是否在相同演示数据下比直接行为克隆提供更强控制能力仍不明确。本文比较了直接行为克隆策略、模仿训练的世界动作策略,以及基于动作条件世界模型优化的策略。在控制器级别,所有世界动作策略均可等效为同闭环轨迹分布的随机策略;在群体层面,在可实现性、精确优化、共用部署信息及分布保持部署条件下,直接行为克隆与世界动作模仿均能恢复观测行为策略。因此,未来预测仅改变学习分解方式,不改变无限制外部策略类或理想模仿目标。动作条件世界模型则通过预测特定动作下的结果并以控制目标进行比较而不同。本文刻画了不依赖候选动作的未来模型中不可消除的动作特异性预测误差,识别出世界-动作联合体恢复干预前向模型的条件,并表明观测演示通常无法识别动作效应。最后构造了一类环境家族,其中所有观测学习者均有正的最坏情况后悔,而一次有效干预即可实现零后悔。核心区别在于:预测观察行为的未来,与预测指定动作的后果用于策略优化。
原文摘要 · Abstract (English)
World-action models predict a future outcome and then infer an associated action. Although this factorization can improve representation learning and data efficiency, it is unclear whether it provides stronger control capability than direct behavior cloning when both are trained from the same observational demonstrations. We compare a direct behavior-cloning policy, an imitation-trained world-action policy, and a policy optimized with an action-conditioned world model. At the controller-class level, every world-action policy can be flattened into a direct stochastic policy with the same closed-loop trajectory distribution. At the population level, under realizability, exact optimization, common deployment information, and distribution-preserving deployment, direct behavior cloning and world-action imitation both recover the observational behavior policy. Thus, future prediction changes the learning factorization but not the unrestricted external policy class or ideal imitation target. Action-conditioned world-model learning differs by predicting outcomes under specified actions and comparing them through a control objective. We characterize the irreducible action-specific prediction error of future models that do not condition on the candidate action, identify conditions under which a world-action joint can recover an interventional forward model, and show that observational demonstrations do not identify action effects in general. Finally, we construct an environment family in which every observational learner has positive worst-case regret, whereas one informative intervention permits zero regret. The key distinction is therefore between predicting futures associated with observed behavior and predicting consequences of specified actions for policy optimization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。