arXiv:2502.09680cs.CVcs.AI2025-02AAAI被引 10

用物体中心方法提升视觉干扰下的动作学习鲁棒性

Object-Centric Latent Action Learning

  • 以物体为中心建模,分离代理与背景运动
  • 在复杂干扰任务中性能提升50%(返利/成功率)
  • 适合需要少量标注数据的智能体高效适应场景

利用海量无标签互联网视频数据进行具身智能是当前瓶颈,主要受限于缺乏动作标签和存在与动作相关的视觉干扰。尽管近期的潜在动作策略优化(LAPO)从视觉观测中推断代理动作标签已展现潜力,但在干扰存在时性能显著下降。为此,我们提出一种新型物体中心潜在动作学习框架,聚焦物体而非像素。通过自监督物体中心预训练,解耦代理运动与背景动态,使LAPO专注于任务相关交互,从而生成更鲁棒的代理动作标签,支持更好的模仿学习和仅需少量标注轨迹即可实现高效智能体适配。我们在八个视觉复杂的任务上评估该方法,涵盖干扰控制套件(DCS)和干扰元世界(DMW)。结果表明,物体中心预训练将干扰负面影响降低50%,以下游任务表现衡量:平均回报(DCS)与成功概率(DMW)。

原文摘要 · Abstract (English)

Leveraging vast amounts of unlabeled internet video data for embodied AI is currently bottlenecked by the lack of action labels and the presence of action-correlated visual distractors. Although recent latent action policy optimization (LAPO) has shown promise in inferring proxy action labels from visual observations, its performance degrades significantly when distractors are present. To address this limitation, we propose a novel object-centric latent action learning framework that centers on objects rather than pixels. We leverage self-supervised object-centric pretraining to disentangle the movement of the agent and distracting background dynamics. This allows LAPO to focus on task-relevant interactions, resulting in more robust proxy-action labels, enabling better imitation learning and efficient adaptation of the agent with just a few action-labeled trajectories. We evaluated our method in eight visually complex tasks across the Distracting Control Suite (DCS) and Distracting MetaWorld (DMW). Our results show that object-centric pretraining mitigates the negative effects of distractors by 50%, as measured by downstream task performance: average return (DCS) and success rate (DMW).

动作学习物体中心具身智能自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。