用汇总统计量计算全局阈值,提升联邦异常检测精度。
Enhanced Federated Anomaly Detection Through Autoencoders Using Summary Statistics-Based Thresholding
- 基于正常与异常数据的汇总统计量,联邦聚合计算全局阈值。
- 在非独立同分布数据下,检测准确率显著优于现有方法。
- 适合需要隐私保护的分布式异常检测场景,如金融风控。
在联邦学习(FL)中,由于数据分散且分布非独立同分布(non-IID),异常检测(AD)面临挑战。本文提出一种新型联邦阈值计算方法,利用正常与异常数据的汇总统计量,提升基于自编码器(AE)的联邦异常检测精度与鲁棒性。该方法在客户端聚合本地统计量,计算出能最优分离异常与正常数据的全局阈值,同时保障隐私。实验在信用卡欺诈检测、Shuttle、Covertype等公开数据集上进行,涵盖多种数据分布场景。结果表明,该方法在处理非IID数据时持续优于现有联邦及本地阈值方法。研究还分析了不同数据分布和客户端数量对性能的影响,验证了基于汇总统计量的阈值计算在提升联邦异常检测系统可扩展性与准确性方面的潜力。
原文摘要 · Abstract (English)
In Federated Learning (FL), anomaly detection (AD) is a challenging task due to the decentralized nature of data and the presence of non-IID data distributions. This study introduces a novel federated threshold calculation method that leverages summary statistics from both normal and anomalous data to improve the accuracy and robustness of anomaly detection using autoencoders (AE) in a federated setting. Our approach aggregates local summary statistics across clients to compute a global threshold that optimally separates anomalies from normal data while ensuring privacy preservation. We conducted extensive experiments using publicly available datasets, including Credit Card Fraud Detection, Shuttle, and Covertype, under various data distribution scenarios. The results demonstrate that our method consistently outperforms existing federated and local threshold calculation techniques, particularly in handling non-IID data distributions. This study also explores the impact of different data distribution scenarios and the number of clients on the performance of federated anomaly detection. Our findings highlight the potential of using summary statistics for threshold calculation in improving the scalability and accuracy of federated anomaly detection systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。