用大模型分析350万帖,发现四分之一用户初帖就明说犯罪。
Assessing Crime Disclosure Patterns in a Large-Scale Cybercrime Forum
- 用三类标签+大模型自动标注,识别用户发帖的犯罪披露程度
- 超三分之一用户初帖提及犯罪,但多数人逐步升级披露
- 灰色内容占比高,说明多数人用模糊表达规避风险,适合执法研究
网络犯罪论坛是网络犯罪生态的核心,为非法商品、服务和知识交换提供平台。以往研究多关注其市场与社交结构,但对用户行为动态,尤其是犯罪活动披露行为了解有限。本研究首次对大型网络犯罪论坛进行大规模评估,分析近30万用户发布的超过350万条帖子。采用三级分类体系(良性、灰色、犯罪)及基于大语言模型(LLMs)的可扩展标注流程,测量初始发帖中的犯罪披露水平,分析用户在不同类别间的转换行为,并探讨犯罪披露与私信通信的关系。结果表明,犯罪披露具有相对普遍性:四分之一的初始帖子包含明确犯罪内容,超过三分之一的用户至少在一次初始发帖中披露犯罪活动。同时,多数参与者保持克制,超过三分之二仅发布良性或灰色内容,且通常逐步升级披露。灰色初始帖子尤为突出,表明许多用户避免直接声明,转而使用模糊表述。研究证明大模型文本分类与马尔可夫链建模在捕捉犯罪披露模式上的价值,为执法部门区分良性、灰色和犯罪内容提供了新思路。
原文摘要 · Abstract (English)
Cybercrime forums play a central role in the cybercrime ecosystem, serving as hubs for the exchange of illicit goods, services, and knowledge. Previous studies have explored the market and social structures of these forums, but less is known about the behavioral dynamics of users, particularly regarding participants' disclosure of criminal activity. This study provides the first large-scale assessment of crime disclosure patterns in a major cybercrime forum, analysing over 3.5 million posts from nearly 300k users. Using a three-level classification scheme (benign, grey, and crime) and a scalable labelling pipeline powered by large language models (LLMs), we measure the level of crime disclosure present in initial posts, analyse how participants switch between levels, and assess how crime disclosure behavior relates to private communications. Our results show that crime disclosure is relatively normative: one quarter of initial posts include explicit crime-related content, and more than one third of users disclose criminal activity at least once in their initial posts. At the same time, most participants show restraint, with over two-thirds posting only benign or grey content and typically escalating disclosure gradually. Grey initial posts are particularly prominent, indicating that many users avoid overt statements and instead anchor their activity in ambiguous content. The study highlights the value of LLM-based text classification and Markov chain modelling for capturing crime disclosure patterns, offering insights for law enforcement efforts aimed at distinguishing benign, grey, and criminal content in cybercrime forums.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。