提出LARY基准,评估视觉到动作的通用表示能力。
LARY: A Latent Action Representation Yielding Benchmark for Generalizable Vision-to-Action Alignment
- 构建统一框架,评估视觉到动作的潜在动作表征。
- 通用视觉模型在无动作监督下表现优于专用模型。
- 语义抽象比像素重建更有效实现视觉到动作对齐。
尽管显式动作数据稀缺限制了视觉-语言-动作(VLA)模型发展,人类动作视频却提供了可扩展且无标签的数据源。利用大规模人类视频数据集的关键挑战在于将视觉信号转化为与本体无关的潜在动作表示。然而,潜在动作表示能否从视觉观察中生成稳健控制尚未得到严谨评估。本文提出潜行动作表示生成(LARY)基准,一个统一框架,用于评估潜在动作表示在高层语义动作(做什么)和低层机器人控制(如何做)上的性能。该全面构建的数据集包含超过一百万段视频(1000小时),涵盖151个动作类别,以及62万张图像对和59.5万条运动轨迹,覆盖多种具身形态和环境。实验揭示两个关键发现:(i) 在无任何动作监督下训练的通用视觉基础模型,持续优于专用具身潜在动作模型;(ii) 基于潜在空间的视觉表示在本质上比基于像素空间更贴近物理动作空间。结果表明,通用视觉表示天然编码了用于物理控制的动作相关知识,且语义级抽象是从视觉到动作的更根本有效路径。
原文摘要 · Abstract (English)
While the shortage of explicit action data limits Vision-Language-Action (VLA) models, human action videos offer a scalable yet unlabeled data source. A critical challenge in utilizing large-scale human video datasets lies in transforming visual signals into ontology-independent representations, known as latent actions. However, the capacity of latent action representation to derive robust control from visual observations has yet to be rigorously evaluated. We introduce the Latent Action Representation Yielding (LARY) Benchmark, a unified framework for evaluating latent action representations on both high-level semantic actions (what to do) and low-level robotic control (how to do). The comprehensively curated dataset encompasses over one million videos (1,000 hours) spanning 151 action categories, alongside 620K image pairs and 595K motion trajectories across diverse embodiments and environments. Our experiments reveal two crucial insights: (i) General visual foundation models, trained without any action supervision, consistently outperform specialized embodied latent action models. (ii) Latent-based visual space is fundamentally better aligned to physical action space than pixel-based space. These results suggest that general visual representations inherently encode action-relevant knowledge for physical control, and that semantic-level abstraction serves as a fundamentally more effective pathway from vision to action than pixel-level reconstruction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。