arXiv:2608.19613cs.ROcs.CV2026-08

系统研究机器人学习中潜在动作的关键设计因素。

What Matters for Latent Actions in Robot Learning

论文配图:What Matters for Latent Actions in Robot Learning
图 1 · 摘自论文原文
  • 统一多种潜在动作模型,构建可比实验框架。
  • 发现微调视觉语言模型能显著提升下游策略性能。
  • 揭示代理指标对真实任务表现的预测有效性。

潜在动作模型(LAM)通过利用大规模未标注视频中的潜在动作,为机器人学习提供了紧凑的物理动作替代表示。尽管进展迅速,现有研究在不同实验设置下孤立评估各类设计选择,缺乏统一比较,难以识别真正影响下游机器人操作性能的因素。本文首次对机器人操作中的潜在动作学习进行系统性实证研究。我们将在统一自编码框架内整合代表性LAM方法,系统考察41种设计选择,涵盖潜在动作建模范式、学习目标与正则化方法、潜在动作融合策略三个维度。进一步评估四种潜在动作质量的代理指标,检验其对下游机器人操作性能的预测能力。在三个广泛使用的基准上的大量实验表明,使用潜在动作微调视觉语言模型(VLM)能为下游策略学习提供更强初始化,且在真实机器人操作任务中得到验证。

原文摘要 · Abstract (English)

Latent Action Models (LAMs) have emerged as a promising paradigm for enabling robot learning to leverage large-scale unlabeled videos through latent actions that serve as compact surrogates for physical actions. Despite rapid progress, research on LAM remains highly fragmented, with existing methods evaluating different design choices in isolation under inconsistent experimental settings, making it difficult to identify the factors that truly determine downstream robotic manipulation performance. In this work, we present the first comprehensive empirical study of latent action learning for robotic manipulation. We unify representative LAM methods within a common autoencoding framework and systematically investigate 41 LAM design choices across three dimensions, including latent action modeling paradigms, learning objectives and regularization methods, and latent action integration strategies. We further examine four proxy metrics for evaluating latent action quality and assess their ability to reliably predict downstream robotic manipulation performance. Extensive experiments on three widely used benchmarks provide strong empirical evidence that fine-tuning vision-language model (VLM) backbones with latent actions provides a stronger initialization for downstream policy learning, with further validation on real-world robot manipulation tasks.

机器人学习潜在动作视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。