arXiv:2607.04184cs.AIcs.LG2026-07中稿 · publication in IEE…

用聚类方法从百万交易中找出可疑行为,识别出超2%异常交易。

A Clustering-Based Framework for Identifying Suspicious Trading Patterns in Capital Market

论文配图:A Clustering-Based Framework for Identifying Suspicious Trading Patterns in Capital Market
图 1 · 摘自论文原文
  • 基于K-Means++聚类,结合市场规则阈值划分交易模式。
  • 检测出2.02%可疑交易,其中51.10%为幌骗行为。
  • 无需真实标签,靠轮廓系数验证效果,适合金融风控参考。

市场操纵是通过人为操控股价以快速获利的不当行为,严重损害交易平台信任度。本文构建了一个无监督欺诈检测工具包,采用K-Means++聚类方法处理2012至2024年间约一百万条金融交易数据。研究提出一种基于聚类的分析流程,利用市场实践中的启发式阈值识别并分类欺诈交易。结果表明,2.02%的交易被标记为可疑,其中51.10%明确指向幌骗(spoofing),0.10%为拉高出货(pump and dump),0.55%涉及内幕交易(insider trading),1.43%为虚假突破(fake breakout),其余46.83%未分类。尽管缺乏真实标签,模型性能通过轮廓系数0.561得到验证。

原文摘要 · Abstract (English)

Market manipulation is the dubious practice of manipulating stock prices in order to make a quick profit, which truly degrades confidence on trading platforms. We implemented an unsupervised fraud-detection toolkit that begins with K-Means++ clustering to address this issue. A dataset of roughly one million financial transactions from 2012 to 2024 is used. In order to identify fraudulent trades and categorize them using market practice heuristic thresholds, the study suggests a clustering-based pipeline. The method highlights 2.02% of trades as suspicious where 51.10% clearly indicate spoofing, 0.10% indicate pump and dump, 0.55% indicate insider trading, 1.43% indicate a fake breakout, and 46.83% are unclassified. Despite the lack of ground truth, the model's performance is confirmed by a Silhouette Score of 0.561.

金融风控异常检测聚类分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。