arXiv:2602.14401cs.CVcs.AI2026-02中稿 · the IEEE INFOCOM 2…

针对智能体视觉语言导航的隐私与异构难题,提出自适应个性化联邦学习框架。

pFedNavi: Structure-Aware Personalized Federated Vision-Language Navigation for Embodied AI

  • 按层动态识别客户端关键组件,细粒度融合参数以兼顾共享与本地化
  • 在R2R和RxR上提升导航成功率7.5%,轨迹保真度提升7.8%,收敛快1.38倍
  • 适合处理设备端数据异构性强的视觉语言导航任务

视觉语言导航(VLN)需要大量来自私有室内环境的轨迹指令数据,引发严重隐私担忧。联邦学习(FL)通过数据本地化缓解此问题,但传统FL在环境与指令风格差异极大的情况下表现不佳,导致单一全局模型效果有限。本文提出pFedNavi,一种结构感知且动态自适应的个性化联邦学习框架,专为VLN设计。核心思想是仅在关键位置个性化:通过逐层混合系数自适应识别客户端特定层,并对选定组件(如编码器-解码器投影层、环境敏感解码器层)进行细粒度参数融合,平衡全局知识共享与局部特化。我们在两个标准VLN基准R2R和RxR上评估,使用ResNet与CLIP视觉表示。所有指标下,pFedNavi均显著优于基于FedAvg的基线,导航成功率达7.5%提升,轨迹保真度提升7.8%,非IID条件下收敛速度加快1.38倍。

原文摘要 · Abstract (English)

Vision-Language Navigation VLN requires large-scale trajectory instruction data from private indoor environments, raising significant privacy concerns. Federated Learning FL mitigates this by keeping data on-device, but vanilla FL struggles under VLNs' extreme cross-client heterogeneity in environments and instruction styles, making a single global model suboptimal. This paper proposes pFedNavi, a structure-aware and dynamically adaptive personalized federated learning framework tailored for VLN. Our key idea is to personalize where it matters: pFedNavi adaptively identifies client-specific layers via layer-wise mixing coefficients, and performs fine-grained parameter fusion on the selected components (e.g., the encoder-decoder projection and environment-sensitive decoder layers) to balance global knowledge sharing with local specialization. We evaluate pFedNavi on two standard VLN benchmarks, R2R and RxR, using both ResNet and CLIP visual representations. Across all metrics, pFedNavi consistently outperforms the FedAvg-based VLN baseline, achieving up to 7.5% improvement in navigation success rate and up to 7.8% gain in trajectory fidelity, while converging 1.38x faster under non-IID conditions.

联邦学习视觉导航个性化多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。