从工业视频中自动提取动作片段,用于机器人预训练。
From Observation to Action: Latent Action-based Primitive Segmentation for VLA Pre-training in Industrial Settings
- 用轻量级运动分词器编码动作动态,再通过潜空间能量度量无监督分割动作原型。
- 在公开数据集和自研电机装配数据上实现关键任务的准确分割。
- 首个端到端自动化系统,适合制造场景下的具身智能研究。
我们提出一种新型无监督框架,利用连续工业视频流中海量未标注的人类示范数据,为视觉-语言-动作(VLA)模型预训练提供支持。方法首先训练一个轻量级运动分词器以编码运动动态,随后采用基于新提出的“潜空间动作能量”度量的无监督动作分割器,发现并分割出语义一致的动作原型。整个流程输出结构化的视频片段及其对应的潜空间动作序列,可直接用于VLA预训练。在公开基准与自有电动机装配数据集上的评估表明,该方法能有效分割工作站中人类执行的关键任务。进一步通过视觉-语言模型进行聚类与定量评估,验证了所发现动作原型的语义一致性。据我们所知,这是首个完全自动化、端到端从非结构化工业视频中提取并组织VLA预训练数据的系统,为制造场景中的具身智能集成提供了可扩展解决方案。
原文摘要 · Abstract (English)
We present a novel unsupervised framework to unlock vast unlabeled human demonstration data from continuous industrial video streams for Vision-Language-Action (VLA) model pre-training. Our method first trains a lightweight motion tokenizer to encode motion dynamics, then employs an unsupervised action segmenter leveraging a novel "Latent Action Energy" metric to discover and segment semantically coherent action primitives. The pipeline outputs both segmented video clips and their corresponding latent action sequences, providing structured data directly suitable for VLA pre-training. Evaluations on public benchmarks and a proprietary electric motor assembly dataset demonstrate effective segmentation of key tasks performed by humans at workstations. Further clustering and quantitative assessment via a Vision-Language Model confirm the semantic coherence of the discovered action primitives. To our knowledge, this is the first fully automated end-to-end system for extracting and organizing VLA pre-training data from unstructured industrial videos, offering a scalable solution for embodied AI integration in manufacturing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。