在复杂场景中,用分解的动态片段构建更精准的隐式动作表示。
Latent Actions from Factorized Transition Effects under Agent Ambiguity

- 先分解观察到的动态为可复用的局部片段,再按状态组合成隐式动作。
- 新方法在动作对齐和下游监督利用上显著优于现有模型。
- 适用于多物体、干扰多的视觉任务,尤其适合无标注场景建模。
隐式动作模型(LAMs)从观测中学习类动作代理。但在多物体或干扰丰富的场景中,观测包含不仅有主体运动,还有干扰物、相机动态和背景变化,导致在无监督条件下难以准确恢复真实动作。我们提出,合适的无监督目标并非真实动作本身,而是场景中存在过渡效应的状态条件组合摘要,从而实现更好的动作对齐并更高效地利用动作监督。为此,我们设计两阶段框架:首先预训练观测过渡分解(OTF),通过组合代码本发现可复用的局部过渡基元;随后将这些基元聚合为紧凑的状态条件隐式动作,分别在标准逆-前向动力学框架下实现为OTF-LAM-Pixel,以及无需解码器的变体OTF-LAM-Dino,后者在冻结的DINOv2表征空间中运行。实验表明,所学过渡基元可在不同视觉外观和形态间迁移,生成的隐式动作表现出更强的动作对齐能力,更有效利用下游动作监督,并在下游策略任务中达到竞争力或更优性能。
原文摘要 · Abstract (English)
Latent Action Models (LAMs) learn action-like proxies from observation. However, in multi-object or distractor-rich scenes, observations contain not only agent motion but also distractors, camera dynamics, and background changes, making recovery of the underlying action intrinsically ambiguous without supervision. We argue that the appropriate unsupervised target is therefore not the true action itself, but a state-conditioned compositional summary of the transition effects present in the scene, enabling better action alignment and more effective utilization of action supervision. To this end, we propose a two-stage framework. We first pretrain Observed Transition Factorization (OTF) to discover reusable local transition primitives using a compositional codebook. We then aggregate these primitives into compact state-conditioned latent actions, instantiated as OTF-LAM-Pixel under the standard inverse-forward dynamics framework and OTF-LAM-Dino, a decoder-free variant operating in a frozen DINOv2 representation space. Experiments show that the learned transition primitives transfer across visual appearance and morphology, while the resulting latent actions exhibit substantially stronger action alignment than existing LAMs, make more effective use of downstream action supervision, and achieve competitive or superior downstream policy performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。