用深度学习自动抓取暗网商品信息,准确率超90%。
Scraping the Shadows: Deep Learning Breakthroughs in Dark Web Intelligence
- 构建新标注数据集,训练三类先进命名实体识别模型。
- 模型在暗网商品页提取中达91%精确率、96%召回率。
- 微调后通用模型表现最佳,适合执法机构情报分析。
暗网市场(DNMs)在全球范围内助长非法商品交易。手动提取其数据费时且易出错。为实现自动化,我们开发了一个从暗网市场提取数据的框架,并评估了三种最先进的命名实体识别(NER)模型——ELMo-BiLSTM、UniversalNER 和 GLiNER——在提取复杂实体任务中的表现。我们提出了一个新的标注数据集,用于模型的训练、微调与评估。结果表明,当前先进的NER模型在暗网商品页面信息提取中表现优异,达到91%精确率、96%召回率和94%的F1分数。此外,微调显著提升性能,其中UniversalNER表现最佳。
原文摘要 · Abstract (English)
Darknet markets (DNMs) facilitate the trade of illegal goods on a global scale. Gathering data on DNMs is critical to ensuring law enforcement agencies can effectively combat crime. Manually extracting data from DNMs is an error-prone and time-consuming task. Aiming to automate this process we develop a framework for extracting data from DNMs and evaluate the application of three state-of-the-art Named Entity Recognition (NER) models, ELMo-BiLSTM \citep{ShahEtAl2022}, UniversalNER \citep{ZhouEtAl2024}, and GLiNER \citep{ZaratianaEtAl2023}, at the task of extracting complex entities from DNM product listing pages. We propose a new annotated dataset, which we use to train, fine-tune, and evaluate the models. Our findings show that state-of-the-art NER models perform well in information extraction from DNMs, achieving 91% Precision, 96% Recall, and an F1 score of 94%. In addition, fine-tuning enhances model performance, with UniversalNER achieving the best performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。