用光流约束提升视频中隐式动作学习的鲁棒性
LAOF: Robust Latent Action Learning with Optical Flow Constraints
- 利用光流作为动作驱动信号,抑制背景干扰
- 在仅1%标签下性能超越监督方法,在10%标签下仍保持优势
- 适合标签稀缺场景下的智能体预训练,尤其适用于机器人任务
从大规模视频中学习隐式动作是构建可扩展具身基础模型预训练的关键,但现有方法常受无关动作干扰影响。尽管引入动作监督可缓解干扰,但其效果受限于动作标签稀缺。光流表征连续帧间的像素级运动,天然抑制背景并突出运动物体。为此,我们提出鲁棒隐式动作学习框架LAOF,通过代理的光流作为动作驱动信号,实现对干扰的鲁棒隐式动作表示学习。实验表明,LAOF在下游模仿学习与强化学习任务中表现优于现有方法。这一优势源于光流约束显著稳定训练,并在极端标签稀疏条件下提升隐式表示质量;即使标签比例增至10%,仍保持有效性。尤为重要的是,无需动作监督时,LAOF性能仍可匹配或超越仅使用1%标签的监督方法。
原文摘要 · Abstract (English)
Learning latent actions from large-scale videos is crucial for the pre-training of scalable embodied foundation models, yet existing methods often struggle with action-irrelevant distractors. Although incorporating action supervision can alleviate these distractions, its effectiveness is restricted by the scarcity of available action labels. Optical flow represents pixel-level motion between consecutive frames, naturally suppressing background elements and emphasizing moving objects. Motivated by this, we propose robust Latent Action learning with Optical Flow constraints, called LAOF, a pseudo-supervised framework that leverages the agent's optical flow as an action-driven signal to learn latent action representations robust to distractors. Experimental results show that the latent representations learned by LAOF outperform existing methods on downstream imitation learning and reinforcement learning tasks. This superior performance arises from optical flow constraints, which substantially stabilize training and improve the quality of latent representations under extremely label-scarce conditions, while remaining effective as the proportion of action labels increases to 10 percent. Importantly, even without action supervision, LAOF matches or surpasses action-supervised methods trained with 1 percent of action labels.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。