提出分层隐动作模型,从无动作视频中发现长期高阶技能。
Hierarchical Latent Action Model
- 用预训练低层模型提取动态模式,分层聚合为高层技能
- 在长时序上显著提升技能发现能力,优于基线模型
- 适合机器人控制与交互世界模型等需要长期规划的任务
隐动作模型(LAMs)可从无动作数据中学习,应用于机器人控制到交互式世界模型。然而,现有LAMs通常只关注短时程帧间转换,捕捉低层运动而忽视更长时间结构。实际上,无动作视频常包含时序扩展的高阶技能。本文提出HiLAM,一种分层隐动作模型,通过建模长时序信息来发现隐含技能。利用预训练的LAM作为底层提取器,将包含视频内在动态模式的隐动作序列聚合为高层隐技能。实验表明,HiLAM优于基线模型,展现出鲁棒的动态技能发现能力。
原文摘要 · Abstract (English)
Latent Action Models (LAMs) enable learning from actionless data for applications ranging from robotic control to interactive world models. However, existing LAMs typically focus on short-horizon frame transitions and capture low-level motion while overlooking longer-term temporal structure. In contrast, actionless videos often contain temporally extended and high-level skills. We present HiLAM, a hierarchical latent action model that discovers latent skills by modeling long-term temporal information. To capture these dependencies across long horizons, we utilize a pretrained LAM as a low-level extractor. This architecture aggregates latent action sequences, which contain the underlying dynamic patterns of the video, into high-level latent skills. Our experiments demonstrate that HiLAM improves over the baseline and exhibits robust dynamic skill discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。