arXiv:2604.03340cs.CVcs.AI2026-04被引 6

让智能体动作更精准:通过加法组合结构学习运动本质

Learning Additively Compositional Latent Actions for Embodied AI

  • 引入加法组合约束,使潜在动作具可叠加、可逆的数学结构
  • 在模拟与真实桌面任务中,动作更聚焦运动本身且量级准确
  • 适合需要精确动作理解的机器人控制场景

潜在动作学习从视觉变化中推断伪动作标签,为利用互联网规模视频提升具身智能提供新路径。然而,多数方法缺乏对物理运动加法性、组合性结构的先验约束,导致潜在表示纠缠无关场景细节或未来信息,错误建模运动幅度。本文提出加法组合潜在动作模型(AC-LAM),在短时空中施加场景级加法组合结构约束。该约束促使潜在动作空间具备身份、逆元、循环一致性等简单代数性质,并抑制非加法组合信息。实验表明,AC-LAM 学习到的潜在动作更具结构性、更专注运动且位移校准良好,在模拟与真实世界桌面任务中均优于当前最优的潜在动作模型。

原文摘要 · Abstract (English)

Latent action learning infers pseudo-action labels from visual transitions, providing an approach to leverage internet-scale video for embodied AI. However, most methods learn latent actions without structural priors that encode the additive, compositional structure of physical motion. As a result, latents often entangle irrelevant scene details or information about future observations with true state changes and miscalibrate motion magnitude. We introduce Additively Compositional Latent Action Model (AC-LAM), which enforces scene-wise additive composition structure over short horizons on the latent action space. These AC constraints encourage simple algebraic structure in the latent action space~(identity, inverse, cycle consistency) and suppress information that does not compose additively. Empirically, AC-LAM learns more structured, motion-specific, and displacement-calibrated latent actions and provides stronger supervision for downstream policy learning, outperforming state-of-the-art LAMs across simulated and real-world tabletop tasks.

具身智能动作学习加法结构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。