arXiv:2605.20223cs.CV2026-05被引 1

揭示隐式动作模型失效根源并提出有效缓解方法

Why Latent Actions Fail, and How to Prevent It

论文配图:Why Latent Actions Fail, and How to Prevent It
图 1 · 摘自论文原文
  • 通过线性框架显式建模外部状态,分析干扰机制
  • 标准重建目标会引入未来外部信息,导致表征偏差
  • 聚焦内生成分的表示空间可有效抑制噪声干扰

隐式动作模型(LAMs)试图从无标签视频中学习类动作表征,通过压缩帧间变化实现。然而,真实视频中的帧不仅包含代理自身状态,还包含背景杂乱等外生状态。这些与动作无关的变化会干扰可靠的隐式动作学习。本文通过扩展线性LAM框架,显式建模外生状态,进行理论分析,发现两个关键点:(1) 最小化标准重建目标会导致隐式动作编码未来观测中的外生信息;(2) 在聚焦内生成分的表示空间中学习,是缓解噪声干扰的关键。进一步证明,以往提出的辅助目标(如动作监督)能有效促使隐式动作在不同外生状态下保持一致性。实验在线性和非线性LAM上验证了上述结论,为外生状态如何阻碍隐式动作学习提供了统一的理论解释,并阐明常见缓解方法的作用机理。

原文摘要 · Abstract (English)

Latent action models (LAMs) aim to learn action-like representations from unlabeled videos by compressing frame-to-frame changes. The frames of in-the-wild videos, however, contain not only the agent's own state but exogenous state such as background clutter. Since the exogenous state introduces changes unrelated to actions, it hinders reliable latent action learning. This paper investigates this problem analytically by extending a linear LAM framework to explicitly model exogenous state. Our analysis reveals two insights: (1) minimizing the standard reconstruction objective produces latent actions that encode exogenous information from future observation; and (2) learning in a representation space that focuses on endogenous components is a key to mitigating the interference of noise. We further show that previously proposed auxiliary objectives, such as action-supervision, provably encourage latent actions to be consistent across exogenous states. These findings are validated through experiments on both linear and nonlinear LAMs, providing a unified theoretical analysis of how exogenous state hinders latent action learning and why common remedies work.

动作表征视频理解自监督学习表征学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。