用Telegram消息实时识别加密货币拉高出货骗局并精准定位目标币种和交易所
PumpSense: Real-Time Detection and Target Extraction of Crypto Pump-and-Dumps on Telegram

- 构建28万条标注数据集,直接分析聊天消息实现秒级检测
- 轻量模型9.4秒/样本,Transformer模型仅50毫秒/样本,精度达F1=0.83
- 首次验证LLM在币种提取中准确率高达91%,优于传统规则方法
通过Telegram协调的加密货币拉高出货骗局威胁市场公平。现有研究因缺乏公开的消息级标注数据及设计局限,难以兼顾检测精度与响应速度。本文提出包含28万余条来自39个操盘群的标注消息的数据集,人工标注出2,246条拉高出货公告及其目标币种与交易所。基于此,定义两个任务:实时公告检测与目标币种/交易所提取。检测方面,对比两种模型:轻量级树模型LightGBM(F1=0.79,延迟9.4秒/样本)与Transformer模型BGE-M3(F1=0.83,延迟50毫秒/样本)。结果表明,消息分析可实现单消息窗口微秒级响应。相比依赖市场数据、通常滞后数十秒发现异常的旧方法,本方案直接分析协调消息,在普通硬件上实现近实时检测。此外,首次建立被操纵币种与交易所提取基准。传统规则方法因代号歧义失效,而大模型(LLM)提取准确率达0.91,表现最佳。
原文摘要 · Abstract (English)
Cryptocurrency pump-and-dump schemes coordinated via Telegram threaten market integrity. However, existing research addressing this specific threat has not yet produced solutions that combine reliable results with fast response. This is in part due to the absence of publicly available, message-level labeled data, as well as design choices. In this paper, we address both issues. In particular, we introduce a corpus of over 280,000 Telegram posts from 39 pump-organizing groups, all manually reviewed to identify 2,246 pump announcements and their targeted cryptocurrency and exchange. Leveraging this dataset, we define two tasks: real-time pump-announcement detection and target cryptocurrency/exchange extraction. For detection, we compare two machine-learning models: a lightweight tree-based LightGBM classifier (F1=0.79, latency=9.4 s/sample) and a transformer-based BGE-M3 (F1=0.83, latency=50 ms/sample). With our proposed approach, we show that message analysis can achieve near-instant pump detection at the level of individual Telegram message windows. Unlike prior work that relies purely on market data and typically detects pumps tens of seconds after abnormal trading activity is observed, our method operates directly on the coordination messages themselves and can be evaluated in microseconds per window on commodity hardware. To our knowledge, we also establish the first benchmark for manipulated coin and exchange extraction. We demonstrate that traditional rule-based extraction methods, widely relied upon in prior literature, are ineffective due to ticker ambiguity. In contrast, LLMs achieve the highest accuracy with a score of 0.91.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。