从视觉图像序列中无监督学习动作模型,避免错误自我强化。
Learning Lifted Action Models from Unsupervised Visual Traces

- 联合学习状态、动作与高层动作模型,无需动作标签。
- 用混合整数规划纠正预测错误,提升一致性。
- 适合需要自动构建动作逻辑的机器人与规划系统。
高效构建能捕捉动作前提与效果的模型,是将AI规划应用于真实场景的关键。以往工作多基于状态或动作序列的高层描述学习此类模型。本文面对更挑战性的设定:仅从状态图像序列中学习高层动作模型,且不提供动作观测。提出一种深度学习框架,联合学习状态预测、动作预测与高层动作模型。引入混合整数线性规划(MILP),用于防止预测坍缩与自增强错误。MILP利用部分轨迹的预测状态、动作和动作模型,求解逻辑一致且最接近原预测的状态、动作与动作模型组合。从中提取伪标签以指导后续训练。多领域实验表明,结合MILP修正可帮助模型跳出局部最优,收敛至全局一致解。
原文摘要 · Abstract (English)
Efficient construction of models capturing the preconditions and effects of actions is essential for applying AI planning in real-world domains. Extensive prior work has explored learning such models from high-level descriptions of state and/or action sequences. In this paper, we tackle a more challenging setting: learning lifted action models from sequences of state images, without action observation. We propose a deep learning framework that jointly learns state prediction, action prediction, and a lifted action model. We also introduce a mixed-integer linear program (MILP) to prevent prediction collapse and self-reinforcing errors among predictions. The MILP takes the predicted states, actions, and action model over a subset of traces and solves for logically consistent states, actions, and action model that are as close as possible to the original predictions. Pseudo-labels extracted from the MILP solution are then used to guide further training. Experiments across multiple domains show that integrating MILP-based correction helps the model escape local optima and converge toward globally consistent solutions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。