构建可配置合成数据集,助力反洗钱系统真实场景测试
AMLgentex: Mobilizing Data-Driven Research to Combat Money Laundering
- 基于真实洗钱特征生成可定制交易数据
- 支持部分可观测、动态演化等复杂场景评估
- 适合反洗钱研究者与金融风控从业者使用
洗钱通过将非法资金转入合法经济体系,助长有组织犯罪。尽管每年有数万亿美元被洗钱,但检测率仍很低,因洗钱者逃避监管,已确认案件稀少,机构仅能观察到全球交易网络的局部信息。由于真实交易数据访问受限,合成数据集对开发和评估检测方法至关重要。然而现有数据集往往忽略部分可观测性、时间动态性、策略行为、标签不确定性、类别不平衡及网络级依赖等问题。我们提出 AMLGentex,一个开源工具套件,用于生成逼真且可配置的交易数据,并评估反洗钱检测方法。AMLGentex 能在模拟真实世界挑战的条件下,系统性评估反洗钱系统。通过发布多个国家特定数据集及实用参数指南,旨在赋能研究人员与从业者,建立协作与进步的共同基础。
原文摘要 · Abstract (English)
Money laundering enables organized crime by moving illicit funds into the legitimate economy. Although trillions of dollars are laundered each year, detection rates remain low because launderers evade oversight, confirmed cases are rare, and institutions see only fragments of the global transaction network. Since access to real transaction data is tightly restricted, synthetic datasets are essential for developing and evaluating detection methods. However, existing datasets fall short: they often neglect partial observability, temporal dynamics, strategic behavior, uncertain labels, class imbalance, and network-level dependencies. We introduce AMLGentex, an open-source suite for generating realistic, configurable transaction data and benchmarking detection methods. AMLGentex enables systematic evaluation of anti-money laundering systems under conditions that mirror real-world challenges. By releasing multiple country-specific datasets and practical parameter guidance, we aim to empower researchers and practitioners and provide a common foundation for collaboration and progress in combating money laundering.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。