arXiv:2605.07474cs.CVcs.AI2026-05

无需语言标注,通过联邦学习训练视觉-语言-动作模型

ForgeVLA: Federated Vision-Language-Action Learning without Language Annotations

论文配图:ForgeVLA: Federated Vision-Language-Action Learning without Language Annotations
图 1 · 摘自论文原文
  • 用实体指令分类器将视觉-动作对转为语言指令,构建完整三元组
  • 在多个基准上超越基线,实现任务可区分的特征表示
  • 适合隐私敏感场景下的机器人智能训练

视觉-语言-动作(VLA)模型在通用机器人智能中潜力巨大,但其规模化受制于高质量标注数据的高昂成本。幸运的是,分布在不同领域的带视觉机器人已产生大量视觉-动作配对数据,可用于更高效地扩展VLA训练。然而,由于各种限制,这些原始数据无法集中聚合,且存在严重异构性。为此,本文提出ForgeVLA,一种从分布式视觉-动作对中训练VLA模型的联邦学习框架,无需中央化数据或人工标注。具体而言,每个客户端配备一个具身指令分类器,将视觉-动作对映射到预定义指令集,恢复缺失的语言模态,形成完整的视觉-语言-动作三元组。除三元组构建外,我们还识别出视觉-语言特征坍塌是此前联邦VLA研究中被忽视的关键挑战。为缓解该问题,ForgeVLA结合客户端对比规划损失与服务器端自适应聚合策略,高效学习任务可区分表征。多基准上的大量实验表明,ForgeVLA显著优于其他基线,消融实验进一步验证了各组件的贡献。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models hold great promise for general-purpose robotic intelligence, yet scaling up such models is severely bottlenecked by the high cost of acquiring annotated training data. Fortunately, vision-equipped robots deployed across various domains already produce abundant vision-action pairs that can be leveraged to scale up VLA training more efficiently. However, these raw data cannot be centrally aggregated due to various constraints and also exhibit severe heterogeneity. To address these challenges, in this paper, we propose ForgeVLA, a federated VLA training framework that learns VLA models from distributed vision-action pairs without centralizing raw data or requiring manual annotations. Specifically, each client in ForgeVLA is equipped with an embodied instruction classifier that maps vision-action pairs to a predefined instruction set, recovering the missing language modality and forming complete vision-language-action triplets. Beyond triplet construction, we also identify vision-language feature collapse as a critical challenge that has been largely overlooked in prior federated VLA research. To mitigate this issue, ForgeVLA combines a client-side contrastive planning loss with a server-side adaptive aggregation strategy to learn task-discriminative representations efficiently. Extensive experiments across multiple benchmarks show that ForgeVLA significantly outperforms other baselines, and ablation studies further validate the contribution of each component.

机器人联邦学习多模态无监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。