解决视频异常检测中多方数据异构导致的语义错位问题
FedVAR: Prototype-Aligned Federated Framework for Video Anomaly Recognition

- 用视觉语言模型构建共享原型,对齐各客户端的异常表征
- 在多个非独立同分布场景下,准确率显著优于现有联邦方法
- 适合工业物联网中分布式视频监控系统的异常识别任务
在工业互联网与信息物理系统时代,联邦学习为视频异常识别提供了去中心化的智能范式。该任务对维持高保真数字孪生和保障关键环境安全至关重要。然而,边缘客户端间固有的数据异构性导致了语义错位问题,即各客户端对“正常”与“异常”事件的学习表征不一致。此问题在细粒度异常类别多样的视频异常识别中尤为突出。现有联邦方法主要针对二分类异常检测,无法解决这一错位问题,阻碍了精细识别。本文提出FedVAR,一种弱监督联邦框架,专为视频异常识别设计。利用视觉语言模型的丰富表征,FedVAR采用基于原型的对齐机制,为所有客户端创建共享语义锚点,重置并对齐其视觉与文本特征空间。该过程强制全网络对“正常性”保持一致表征,直接缓解语义错位,并以极低通信开销实现鲁棒的提示学习异常方向向量。我们在多种非独立同分布数据划分、未见领域及新异常类别设置下,在挑战性基准上进行了广泛实验。结果表明,FedVAR始终优于当前最优联邦基线,为基于视频的工控系统分布式智能提供了稳健框架。
原文摘要 · Abstract (English)
In the era of Industrial Internet of Things (IIoT) and Cyber-Physical Systems (CPS), Federated Learning (FL) offers a promising decentralized intelligence paradigm for Video Anomaly Recognition (VAR). This task is vital for maintaining high-fidelity Digital Twins and ensuring safety in mission-critical environments. However, the inherent data heterogeneity across distributed edge clients leads to a fundamental challenge known as semantic misalignment, where clients learn divergent feature representations of "normal" and "abnormal" events. The problem becomes particularly pronounced in VAR, where the presence of diverse and fine-grained anomaly categories leads each client to develop distinct semantic interpretations of abnormality. Existing federated methods primarily focus on binary anomaly detection and fail to address this misalignment, preventing effective fine-grained recognition. In this paper, we introduce FedVAR, a weakly-supervised FL framework explicitly designed for VAR. Leveraging the rich representations of Vision-Language Models (VLMs), FedVAR employs a prototype-based alignment mechanism that creates a shared semantic anchor for all clients to re-center and align their visual and textual feature spaces. This process enforces a consistent representation of "normality" across the decentralized network, directly mitigating semantic misalignment and enabling robust prompt-learning of anomaly direction vectors with minimal communication overhead. We conduct extensive experiments on challenging benchmarks under various non-IID data partitioning schemes, unseen domains, and novel anomaly classes. The results demonstrate that FedVAR consistently outperforms state-of-the-art federated baselines, establishing a robust framework for distributed intelligence in video-based CPS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。