揭秘隐动作模型到底学到了动作还是噪声。
What Do Latent Action Models Actually Learn?
- 构建线性模型分析隐动作学习机制。
- 发现隐变量更易捕捉可控变化而非噪声。
- 适合关注自监督视频表征的研究者。
隐动作模型(LAMs)旨在通过压缩帧间变化的潜在表示,从无标签视频中学习与动作相关的变动。然而,帧间差异可能源于可控变化或外部噪声,引发关键问题:潜在表示是否捕捉了动作引起的变动,还是无关噪声?本文通过可解析的线性模型,揭示了LAM学习的本质,建立了其与主成分分析(PCA)的联系,明确了数据生成策略的理想特性,并验证了数据增强、数据清洗及辅助动作预测等策略在促进学习可控变化方面的合理性。基于数值模拟的可视化结果进一步阐明了观测数据、动作和噪声结构对LAM学习的影响。
原文摘要 · Abstract (English)
Latent action models (LAMs) aim to learn action-relevant changes from unlabeled videos by compressing changes between frames as latents. However, differences between video frames can be caused by controllable changes as well as exogenous noise, leading to an important concern -- do latents capture the changes caused by actions or irrelevant noise? This paper studies this issue analytically, presenting a linear model that encapsulates the essence of LAM learning, while being tractable.This provides several insights, including connections between LAM and principal component analysis (PCA), desiderata of the data-generating policy, and justification of strategies to encourage learning controllable changes using data augmentation, data cleaning, and auxiliary action-prediction. We also provide illustrative results based on numerical simulation, shedding light on the specific structure of observations, actions, and noise in data that influence LAM learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。