统一网络安全命名实体识别标签,提升多数据集泛化能力
Label Unification for Cross-Dataset Generalization in Cybersecurity NER
- 对四个网络安全数据集进行粗粒度标签统一
- 统一训练模型跨数据集表现差,泛化能力有限
- 提出图结构迁移模型,但性能未显著优于基线
网络安全命名实体识别领域缺乏标准化标签,导致数据集难以融合。本文针对四个网络安全数据集开展标签统一研究,实施粗粒度标签对齐,并使用BiLSTM模型进行两两跨数据集评估。定性分析揭示预测错误、模型局限及数据集差异。为克服统一限制,提出多头共享权重模型与基于图的迁移模型。结果表明,统一训练的模型在跨数据集上泛化效果不佳;多头模型仅带来微弱改进;基于BERT-base-NER的图结构迁移模型亦未显著优于原模型。
原文摘要 · Abstract (English)
The field of cybersecurity NER lacks standardized labels, making it challenging to combine datasets. We investigate label unification across four cybersecurity datasets to increase data resource usability. We perform a coarse-grained label unification and conduct pairwise cross-dataset evaluations using BiLSTM models. Qualitative analysis of predictions reveals errors, limitations, and dataset differences. To address unification limitations, we propose alternative architectures including a multihead model and a graph-based transfer model. Results show that models trained on unified datasets generalize poorly across datasets. The multihead model with weight sharing provides only marginal improvements over unified training, while our graph-based transfer model built on BERT-base-NER shows no significant performance gains compared BERT-base-NER.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。