用智能数据提炼技术,让视觉语言动作模型训练更快更高效
FT-NCFM: An Influence-Aware Data Distillation Framework for Efficient VLA Models
- 通过因果归因+程序化对比验证评估数据价值
- 仅用5%精选数据即可达90%任务成功率,训练时间减少80%以上
- 适合追求高效训练的VLA模型研究者与工业应用团队
视觉语言动作(VLA)模型的强大泛化能力受限于其对海量、冗余且价值不均数据集的高度依赖,阻碍了广泛应用。现有以模型为中心的优化路径,如模型压缩(常导致性能下降)或策略蒸馏(产物依赖模型且缺乏通用性),未能从根本上解决数据层面的挑战。为此,本文提出一种全新的、以数据为中心的生成式数据蒸馏框架FT-NCFM。该框架采用自包含的因果溯源(FT)引擎,结合因果归因与程序化对比验证,评估样本的内在价值。基于这些评估,对抗式NCFM过程生成模型无关、信息密集且可复用的数据资产。在多个主流VLA基准上的实验表明,仅用我们提炼出的5%核心数据集进行训练,模型成功率可达85%-90%,同时训练时间减少超过80%。本工作证明,智能数据蒸馏是构建高效高性能VLA模型的一条极具前景的新路径。
原文摘要 · Abstract (English)
The powerful generalization of Vision-Language-Action (VLA) models is bottlenecked by their heavy reliance on massive, redundant, and unevenly valued datasets, hindering their widespread application. Existing model-centric optimization paths, such as model compression (which often leads to performance degradation) or policy distillation (whose products are model-dependent and lack generality), fail to fundamentally address this data-level challenge. To this end, this paper introduces FT-NCFM, a fundamentally different, data-centric generative data distillation framework. Our framework employs a self-contained Fact-Tracing (FT) engine that combines causal attribution with programmatic contrastive verification to assess the intrinsic value of samples. Guided by these assessments, an adversarial NCFM process synthesizes a model-agnostic, information-dense, and reusable data asset. Experimental results on several mainstream VLA benchmarks show that models trained on just 5% of our distilled coreset achieve a success rate of 85-90% compared with training on the full dataset, while reducing training time by over 80%. Our work demonstrates that intelligent data distillation is a highly promising new path for building efficient, high-performance VLA models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。