通过识别可复用行为结构,用一半数据达到全量训练效果
SIEVE: Structure-Aware Data Selection for Imitation Learning with VLA Models

- 将演示分解为视觉-动作基元和过渡接口,捕捉长时序行为结构
- 仅用50%数据和50%训练步数,性能超越全量数据训练
- 适合追求高效训练的机器人模仿学习研究者
视觉-语言-动作(VLA)模型通常在大规模机器人示范数据集上通过模仿学习进行训练,但更多数据并不一定带来更好策略,原因在于数据冗余、噪声以及覆盖不均。现有数据选择方法通常在轨迹或状态-动作层面评估示范,忽略了构成长时序行为的可复用结构。本文提出SIEVE,一种面向VLA模仿学习的结构感知数据选择方法。SIEVE将示范视为可复用基元与转换接口的组合,先从分割后的轨迹中发现视觉-动作基元,再根据收益递减规律,最大化结构重用感知的暴露度来分配选择预算。最后,在每个组合模式桶内选取中位轨迹,保留中心、稳定且适合模仿的示范。在多个数据集、基准和VLA模型上的实验表明,SIEVE始终优于现有基线。值得注意的是,使用仅50%的示范数据和50%的训练步数,SIEVE即可超越全量数据训练,说明通过基元与过渡所捕捉的可复用结构,是高效VLA模仿学习的重要信号。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models are typically trained by imitation learning on large-scale robot demonstration datasets, but more data does not necessarily yield better policies due to redundancy, noise, and uneven coverage. Existing data selection methods often assess demonstrations at either the trajectory or state-action level, missing the reusable structures that compose long-horizon behaviors. In this paper, we propose SIEVE, a structure-aware data selection method for VLA imitation learning. SIEVE views demonstrations as compositions of reusable primitives and transition interfaces. It first discovers visuo-motor primitives from segmented trajectories, then allocates selection budgets to composition patterns by maximizing reuse-aware structural exposure under diminishing returns. Finally, it selects medoid trajectories within each composition-pattern bucket to retain central, stable, and imitation-friendly demonstrations. Experiments across multiple datasets, benchmarks, and VLA models show that SIEVE consistently outperforms competitive data selection baselines. Notably, SIEVE can surpass full-data training while using only 50% of demonstrations and 50% of training steps, suggesting that reusable structure, captured through primitives and transitions, is an important signal for efficient VLA imitation learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。