arXiv:2511.01583cs.CRcs.AI2025-11被引 2

用联邦学习实现跨组织隐私保护的勒索软件检测

Federated Cyber Defense: Privacy-Preserving Ransomware Detection Across Distributed Systems

  • 各机构本地训练模型,通过联邦学习聚合共享特征
  • 相比本地模型,检测准确率提升9%,接近中心化训练效果
  • 适合需跨企业协作又受限于隐私法规的网络安全场景

检测勒索软件对保障云存储、企业文件共享和数据库服务等互联生态安全至关重要。训练高性能人工智能检测模型需要多样化的数据,但这些数据通常分散在多个组织中,集中处理面临安全、隐私法规、数据所有权及法律障碍。此外,勒索软件快速演进,要求模型兼具鲁棒性与可适应性。本文基于Sherpa.ai FL平台评估联邦学习(FL),使多机构可在不交换原始数据的前提下协同训练勒索软件检测模型,该范式特别适用于部署于数百万终端的安全公司(含软硬件厂商)。在严格安全或监管限制下,数据无法离开客户设备。尽管FL可广泛应用于各类恶意软件,本文使用Ransomware Storage Access Patterns(RanSAP)数据集验证其有效性。实验表明,相比服务器本地模型,联邦学习使勒索软件检测准确率相对提升9%,性能接近集中式训练。结果证明,联邦学习为跨组织、跨监管边界提供了一种可扩展、高性能且隐私保护的主动勒索软件检测框架。

原文摘要 · Abstract (English)

Detecting malware, especially ransomware, is essential to securing today's interconnected ecosystems, including cloud storage, enterprise file-sharing, and database services. Training high-performing artificial intelligence (AI) detectors requires diverse datasets, which are often distributed across multiple organizations, making centralization necessary. However, centralized learning is often impractical due to security, privacy regulations, data ownership issues, and legal barriers to cross-organizational sharing. Compounding this challenge, ransomware evolves rapidly, demanding models that are both robust and adaptable. In this paper, we evaluate Federated Learning (FL) using the Sherpa.ai FL platform, which enables multiple organizations to collaboratively train a ransomware detection model while keeping raw data local and secure. This paradigm is particularly relevant for cybersecurity companies (including both software and hardware vendors) that deploy ransomware detection or firewall systems across millions of endpoints. In such environments, data cannot be transferred outside the customer's device due to strict security, privacy, or regulatory constraints. Although FL applies broadly to malware threats, we validate the approach using the Ransomware Storage Access Patterns (RanSAP) dataset. Our experiments demonstrate that FL improves ransomware detection accuracy by a relative 9% over server-local models and achieves performance comparable to centralized training. These results indicate that FL offers a scalable, high-performing, and privacy-preserving framework for proactive ransomware detection across organizational and regulatory boundaries.

联邦学习勒索软件检测隐私保护网络安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。