AnonShield实现漏洞数据快速匿名化,10分钟处理550MB数据。
AnonShield: Scalable On-Premise Pseudonymization for CSIRT Vulnerability Data

- 结合GPU加速命名实体识别与流式处理,实现高吞吐匿名化。
- 处理时间从92小时缩短至10分钟内,达738倍提速,准确率超94%。
- 适合需要合规共享漏洞数据的CSIRT团队使用。
我们提出AnonShield,一种高吞吐、本地部署的匿名化系统,融合GPU加速命名实体识别(NER)、流式处理、缓存机制和面向模式的配置。在最大550 MB(70,951条记录)的数据集上评估,处理时间从超过92小时缩短至10分钟以内(最高738倍加速),同时达到94.2%的F1分数和96.7%的召回率。结果表明,在不牺牲分析效用的前提下,可实现漏洞数据的可扩展匿名化,使运营中的计算机安全事件响应团队(CSIRT)能够合规共享数据。
原文摘要 · Abstract (English)
We present AnonShield, a high-throughput, on-premise pseudonymization system that combines GPU-accelerated NER, streaming processing, caching, and schema-aware configuration. Evaluated on datasets up to 550 MB (70,951 records), AnonShield reduces processing time from over 92 hours to under 10 minutes (up to 738x speedup) while achieving up to 94.2% F1-score and 96.7% recall. Our results show that scalable pseudonymization of vulnerability data is feasible without sacrificing analytical utility, enabling compliant data sharing in operational CSIRT environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。