用分布化隐动作建模时序变化,提升机器人控制的连贯性与泛化能力
DLAM: Distributional Latent Actions with Temporal Constraints

- 将每段动作建模为对角高斯分布,通过均值和方差约束动态一致性
- 在未见视频上重建效果优于基线,累积重建误差降低23%
- 适合做少样本机器人策略迁移,尤其适用于真实场景操控任务
视觉-语言-动作(VLA)模型受限于稀缺的动作标注机器人数据,而无动作视频则提供了丰富的物理变化观测。隐动作模型可提取此类先验,但仅通过重建训练的编码可能无法生成与机器人动作协同的结构化未来。现有结构化方法虽加入时序约束,但保留确定性转移点,局部推断误差会在递归组合中传播放大。本文提出DLAM,一种分布化隐动作模型,将每个转移表示为对角高斯分布。以参考帧为条件的重建使均值基于观测到的视觉变化,而等距三元组上的归一化组合与反转则同时约束均值和维度独立方差。方差组合使用轻量级共享相关系数,捕捉共享中间帧的相邻转移间的依赖关系;反转操作反向均值并保留方差。下游策略学习中,冻结编码器,并训练流匹配策略联合生成均值转移序列与机器人动作。在未见转移上,DLAM学习到比现有基线更时序一致的隐动态,且在未见视频上实现更强的直接与累积重建。在相同的受控π₀迁移协议下,其在MetaWorld MT50、LIBERO及真实世界操控任务中均提升策略性能。受控消融实验表明,归一化均值约束贡献了大部分重建提升,而学习到的方差与相关性感知组合带来互补的下游控制改进。
原文摘要 · Abstract (English)
Vision-language-action (VLA) models remain constrained by scarce action-labeled robot data, whereas action-free videos offer abundant observations of physical change. Latent action models can extract such priors, but reconstruction-trained codes may predict future observations without the structure required for joint generation with robot actions. Existing structured methods add temporal constraints but retain deterministic transition points, so residual errors in locally inferred transitions may propagate and compound under recursive composition. We introduce DLAM, a distributional latent-action model that represents each transition as a diagonal Gaussian. Reconstruction conditioned on the reference frame grounds the mean in observed visual change, while normalized composition and reversal over equal-gap triplets constrain both the mean and dimension-wise variance. Variance composition uses a lightweight shared-correlation coefficient to account for dependence between adjacent transitions that share an intermediate frame, whereas reversal negates the mean and preserves the variance. For downstream policy learning, we freeze the encoder and train a flow-matching policy to jointly generate mean transition sequences and robot actions. On held-out transitions, DLAM learns more temporally consistent latent dynamics than existing latent-action baselines and achieves stronger direct and cumulative reconstruction on held-out videos. Under the same controlled $π_0$ transfer protocol, it also improves policy performance on MetaWorld MT50, LIBERO, and real-world manipulation tasks. Controlled ablations show that normalized mean constraints account for most of the reconstruction gain, while learned variance and correlation-aware composition provide complementary improvements in downstream control.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。