arXiv:2605.10819cs.ROcs.AI2026-05

用视频自动生成结构化动作潜空间,提升机器人任务成功率。

ALAM: Algebraically Consistent Latent Action Model for Vision-Language-Action Models

论文配图:ALAM: Algebraically Consistent Latent Action Model for Vision-Language-Action Models
图 1 · 摘自论文原文
  • 通过视频帧三元组学习具有代数一致性的潜动作转移
  • 在真实任务中将成功率从47.9%提升至85.0%
  • 适合需要长序列动作生成的机器人视觉-语言-动作模型

视觉-语言-动作(VLA)模型受限于带动作标签的机器人数据稀缺,而无动作视频则提供了丰富的物理世界变化证据。潜动作模型可从中提取先验知识,但重建训练的潜码未必适合策略生成:可能仅预测未来观测,缺乏与机器人动作协同生成所需的结构。我们提出ALAM(代数一致潜动作模型),将无动作视频中的时序关系转化为结构化监督。给定帧三元组,ALAM学习由重建约束并受组合与逆向一致性正则化的潜转移,鼓励局部可加的转移空间。下游VLA学习中,冻结预训练编码器,使用其潜转移序列作为辅助生成目标,与机器人动作联合在流匹配目标下生成。这将结构化潜转移与基于流的策略生成耦合,使策略能利用ALAM的局部一致性转移几何,无需潜码到动作的解码。表征探测显示,相比无结构基线,ALAM将可加性与可逆性误差降低25-85倍,并提升长时序累积重建效果。迁移至VLA策略后,其在MetaWorld MT50上平均成功率从47.9%提升至85.0%,在LIBERO上从94.1%提升至98.1%,且在真实操作任务中持续增益。消融实验进一步确认,最强提升源于代数结构潜转移与联合流匹配的协同作用。

原文摘要 · Abstract (English)

Vision-language-action (VLA) models remain constrained by the scarcity of action-labeled robot data, whereas action-free videos provide abundant evidence of how the physical world changes. Latent action models offer a promising way to extract such priors from videos, but reconstruction-trained latent codes are not necessarily suitable for policy generation: they may predict future observations while lacking the structure needed to be reused or generated coherently with robot actions. We introduce ALAM (Algebraic Latent Action Model), an Algebraically Consistent Latent Action Model that turns temporal relations in action-free video into structural supervision. Given frame triplets, ALAM learns latent transitions that are grounded by reconstruction while being regularized by composition and reversal consistency, encouraging a locally additive transition space. For downstream VLA learning, we freeze the pretrained encoder and use its latent transition sequences as auxiliary generative targets, co-generated with robot actions under a joint flow-matching objective. This couples structured latent transitions with flow-based policy generation, allowing the policy to exploit ALAM's locally consistent transition geometry without requiring latent-to-action decoding. Representation probes show that ALAM reduces additivity and reversibility errors by 25-85 times over unstructured latent-action baselines and improves long-horizon cumulative reconstruction. When transferred to VLA policies, ALAM raises the average success rate from 47.9% to 85.0% on MetaWorld MT50 and from 94.1% to 98.1% on LIBERO, with consistent gains on real-world manipulation tasks. Ablations further confirm that the strongest improvements arise from the synergy between algebraically structured latent transitions and joint flow matching.

视觉-语言-动作潜动作模型机器人学习结构化生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。