arXiv:2603.21647cs.CVcs.LG2026-03

解决多视角视频联邦学习中的视角差异与通信开销问题

FedCVU: Federated Learning for Cross-View Video Understanding

  • 引入视图归一化与对比正则化,提升跨视角表征对齐
  • 在动作识别和行人重识别任务中,未见视角准确率显著提升
  • 适合大规模分布式视频分析场景,尤其注重隐私保护

联邦学习(FL)为隐私保护下的多摄像头视频理解提供了新范式。然而,在跨视角场景中面临三大挑战:(i) 不同视角和背景导致客户端数据高度非独立同分布,易过拟合于视角特有模式;(ii) 局部分布偏移引发表征错位,阻碍一致的跨视角语义表达;(iii) 大型视频模型带来巨大通信开销。为此,我们提出 FedCVU 框架,包含三个组件:VS-Norm 保留归一化参数以应对视角特有统计;CV-Align 是轻量级对比正则化模块,增强跨视角表示对齐;SLA 为选择性层聚合策略,降低通信成本且不牺牲精度。在跨视角协议下的动作理解与行人重识别任务中,大量实验表明,FedCVU 在提升未见视角性能的同时保持良好已见视角表现,优于当前最优联邦学习基线,对域异质性和通信约束具有鲁棒性。

原文摘要 · Abstract (English)

Federated learning (FL) has emerged as a promising paradigm for privacy-preserving multi-camera video understanding. However, applying FL to cross-view scenarios faces three major challenges: (i) heterogeneous viewpoints and backgrounds lead to highly non-IID client distributions and overfitting to view-specific patterns, (ii) local distribution biases cause misaligned representations that hinder consistent cross-view semantics, and (iii) large video architectures incur prohibitive communication overhead. To address these issues, we propose FedCVU, a federated framework with three components: VS-Norm, which preserves normalization parameters to handle view-specific statistics; CV-Align, a lightweight contrastive regularization module to improve cross-view representation alignment; and SLA, a selective layer aggregation strategy that reduces communication without sacrificing accuracy. Extensive experiments on action understanding and person re-identification tasks under a cross-view protocol demonstrate that FedCVU consistently boosts unseen-view accuracy while maintaining strong seen-view performance, outperforming state-of-the-art FL baselines and showing robustness to domain heterogeneity and communication constraints.

联邦学习视频理解跨视角隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。