arXiv:2512.11584cs.LGcs.AI2025-12

将复杂操作分解为原子动作,提升通用视觉语言动作模型的泛化能力。

Atomic Action Slicing: Planner-Aligned Options for Generalist VLA Agents

  • 将长程示范拆分为短时序、带类型标签的原子动作,适配规划器使用。
  • 在LIBERO数据集上,任务成功率提升至95.3%(原94.2%)和88.8%(原83.8%)。
  • 适合需要强泛化与组合新技能的机器人控制研究者使用。

当前视觉-语言-动作(VLA)模型泛化能力差,尤其在需要新技能或物体组合的任务中表现不佳。本文提出原子动作切分(AAS),一种与规划器对齐的方法,将长程示范分解为短时序、带类型标签的原子动作,更易被规划器使用且利于策略学习。基于LIBERO示范数据,AAS构建了包含2,124个原子片段的标注数据集,每个片段标注动作类型、时间跨度与置信度。更强的分割器(Gemini 2.5 Pro)能紧密匹配规划定义的计划,并在关键帧抖动下保持鲁棒性;而小型模型在多物体任务中表现较差。在原子数据集上微调CLIP-RT+后,LIBERO-Goal任务成功率从94.2%提升至95.3%,LIBERO-Long任务从83.8%提升至88.8%。数据集已公开发布于HuggingFace(https://huggingface.co/datasets/gate-institute/GATE-VLAP-datasets)。

原文摘要 · Abstract (English)

Current vision-language-action (VLA) models generalize poorly, particularly when tasks require new compositions of skills or objects. We introduce Atomic Action Slicing (AAS), a planner-aligned approach that decomposes long-horizon demonstrations into short, typed atomic actions that are easier for planners to use and policies to learn. Using LIBERO demonstrations, AAS produces a validated dataset of 2,124 atomic segments labeled with action type, temporal span, and confidence. A stronger segmenter (Gemini 2.5 Pro) closely matches planner-defined plans and remains robust under keyframe jitter, while smaller models perform worse on multi-object tasks. Fine-tuning CLIP-RT+ on our atomic dataset improves task success from 94.2% to 95.3% on LIBERO-Goal and 83.8% to 88.8% on LIBERO-Long. We publicly release the GATE-VLAP dataset on HuggingFace(https://huggingface.co/datasets/gate-institute/GATE-VLAP-datasets)

机器人控制动作规划视觉语言动作数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。