从Telegram自动化构建大规模高精度恶意威胁情报数据集
CTI Dataset Construction from Telegram
- 设计端到端流水线,自动抓取并筛选Telegram中的威胁信息
- 从14.5万条消息中提取8.65万条真实恶意指标,分类准确率达96.64%
- 为安全研究与实战检测提供可复用的高质量威胁数据基础
网络威胁情报(CTI)助力组织预判、发现并应对不断演化的网络攻击。其有效性依赖高质量数据集,用于模型开发、训练、评估与基准测试。随着攻击手法持续演变,构建此类数据集至关重要。近期,Telegram因其及时且多样化的威胁信息,成为有价值的CTI来源。本文提出一种端到端自动化流程,系统性地从Telegram收集并过滤威胁相关内容。该流程识别出150个来源,从中筛选12个优质频道,共抓取145,349条消息。为精准区分威胁情报与普通内容,采用基于BERT的分类器,准确率达到96.64%。最终从过滤后的消息中构建包含86,509条恶意指标的数据集,涵盖域名、IP、URL、哈希及CVE等。该方法不仅生成大规模、高保真度的CTI数据集,也为未来网络安全检测的研究与实际应用奠定基础。
原文摘要 · Abstract (English)
Cyber Threat Intelligence (CTI) enables organizations to anticipate, detect, and mitigate evolving cyber threats. Its effectiveness depends on high-quality datasets, which support model development, training, evaluation, and benchmarking. Building such datasets is crucial, as attack vectors and adversary tactics continually evolve. Recently, Telegram has gained prominence as a valuable CTI source, offering timely and diverse threat-related information that can help address these challenges. In this work, we address these challenges by presenting an end-to-end automated pipeline that systematically collects and filters threat-related content from Telegram. The pipeline identifies relevant Telegram channels and scrapes 145,349 messages from 12 curated channels out of 150 identified sources. To accurately filter threat intelligence messages from generic content, we employ a BERT-based classifier, achieving an accuracy of 96.64%. From the filtered messages, we compile a dataset of 86,509 malicious Indicators of Compromise, including domains, IPs, URLs, hashes, and CVEs. This approach not only produces a large-scale, high-fidelity CTI dataset but also establishes a foundation for future research and operational applications in cyber threat detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。