让机器人学指令操作,还能保护用户数据隐私。
FedVLA: Federated Vision-Language-Action Learning with Dual Gating Mixture-of-Experts for Robotic Manipulation
- 用双门控专家混合模型提升多任务理解能力
- 分布式训练下成功率接近中心化方法
- 适合注重隐私的智能机器人研发团队
视觉-语言-动作(VLA)模型显著推动了机器人操作的发展,使机器人能根据语言指令执行任务。然而,训练这些模型通常依赖大规模用户专属数据,引发隐私与安全担忧,限制了其广泛应用。为此,我们提出FedVLA,首个联邦VLA学习框架,实现分布式训练,在不牺牲性能的前提下保护数据隐私。该框架融合任务感知表征学习、自适应专家选择和专家驱动的联邦聚合,实现高效且私密的VLA模型训练。具体而言,我们设计了指令导向场景解析机制,基于任务指令分解并增强物体级特征,提升上下文理解能力;提出双门控专家混合(DGMoE)机制,输入令牌与自知专家共同决定激活状态,有效学习多样化任务模式;进一步在联邦服务器端设计专家驱动聚合策略,由激活专家引导模型聚合,保障跨客户端知识迁移效果。大量仿真与真实机器人实验表明,所提方案有效:相比基线,DGMoE显著提升计算效率,而FedVLA在任务成功率上达到与集中式训练相当水平,充分保护数据隐私。
原文摘要 · Abstract (English)
Vision-language-action (VLA) models have significantly advanced robotic manipulation by enabling robots to interpret language instructions for task execution. However, training these models often relies on large-scale user-specific data, raising concerns about privacy and security, which in turn limits their broader adoption. To address this, we propose FedVLA, the first federated VLA learning framework, enabling distributed model training that preserves data privacy without compromising performance. Our framework integrates task-aware representation learning, adaptive expert selection, and expert-driven federated aggregation, enabling efficient and privacy-preserving training of VLA models. Specifically, we introduce an Instruction Oriented Scene-Parsing mechanism, which decomposes and enhances object-level features based on task instructions, improving contextual understanding. To effectively learn diverse task patterns, we design a Dual Gating Mixture-of-Experts (DGMoE) mechanism, where not only input tokens but also self-aware experts adaptively decide their activation. Finally, we propose an Expert-Driven Aggregation strategy at the federated server, where model aggregation is guided by activated experts, ensuring effective cross-client knowledge transfer.Extensive simulations and real-world robotic experiments demonstrate the effectiveness of our proposals. Notably, DGMoE significantly improves computational efficiency compared to its vanilla counterpart, while FedVLA achieves task success rates comparable to centralized training, effectively preserving data privacy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。