用真实公开数据混合合成数据,提升反洗钱模型效果。
Hybrid Data can Enhance the Utility of Synthetic Data for Training Anti-Money Laundering Models
- 将真实公开特征融入合成数据,构建混合数据集
- 混合数据集使模型准确率显著优于纯合成数据
- 适合需兼顾隐私与模型性能的金融机构使用
洗钱是金融机构面临的重要全球性问题。自动化反洗钱(AML)模型,如图神经网络(GNN),可用于实时识别非法交易。但开发此类模型的一大障碍是因隐私和保密性顾虑而难以获取训练数据。已有研究提出通过生成模拟真实数据统计特性的合成数据来解决此问题,同时保护隐私。然而,仅用纯合成数据训练AML模型仍存在挑战。本文提出使用混合数据集,通过引入公开可得、易于获取的真实世界特征,增强合成数据的实用性。实验表明,混合数据集不仅保持了隐私保护,还显著提升了模型性能,为金融机构优化AML系统提供了可行路径。
原文摘要 · Abstract (English)
Money laundering is a critical global issue for financial institutions. Automated Anti-money laundering (AML) models, like Graph Neural Networks (GNN), can be trained to identify illicit transactions in real time. A major issue for developing such models is the lack of access to training data due to privacy and confidentiality concerns. Synthetically generated data that mimics the statistical properties of real data but preserves privacy and confidentiality has been proposed as a solution. However, training AML models on purely synthetic datasets presents its own set of challenges. This article proposes the use of hybrid datasets to augment the utility of synthetic datasets by incorporating publicly available, easily accessible, and real-world features. These additions demonstrate that hybrid datasets not only preserve privacy but also improve model utility, offering a practical pathway for financial institutions to enhance AML systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。