arXiv:2412.14730cs.LG2024-12被引 5

为银行生成金融交易数据,评估五种生成模型优劣。

Generative AI for Banks: Benchmarks and Algorithms for Synthetic Financial Transaction Data

  • 对比五种生成模型在数据真实度、隐私保护等方面表现。
  • CTGAN综合表现最佳,适合一般场景;DGAN最适配高隐私需求。
  • 研究结果可指导银行选型,提升数据安全与模型训练效率。

银行业因数据敏感性和监管限制难以应用深度学习,生成式AI可能提供解决方案。本研究评估了五种领先模型——条件表格式生成对抗网络(CTGAN)、DoppelGANger(DGAN)、Wasserstein GAN、金融扩散模型(FinDiff)和表格式变分自编码器(TVAE)——在五个维度上的表现:数据保真度、合成质量、效率、隐私保护和图结构保留。结果显示,各模型均无法完全复现真实数据的图结构,但各有专长:DGAN在隐私敏感任务中表现最优,FinDiff与TVAE在数据复制与增强方面突出,而CTGAN在五项指标间取得平衡,适合中等隐私要求的通用场景。研究为金融机构选择合适算法提供了实证依据。

原文摘要 · Abstract (English)

The banking sector faces challenges in using deep learning due to data sensitivity and regulatory constraints, but generative AI may offer a solution. Thus, this study identifies effective algorithms for generating synthetic financial transaction data and evaluates five leading models - Conditional Tabular Generative Adversarial Networks (CTGAN), DoppelGANger (DGAN), Wasserstein GAN, Financial Diffusion (FinDiff), and Tabular Variational AutoEncoders (TVAE) - across five criteria: fidelity, synthesis quality, efficiency, privacy, and graph structure. While none of the algorithms is able to replicate the real data's graph structure, each excels in specific areas: DGAN is ideal for privacy-sensitive tasks, FinDiff and TVAE excel in data replication and augmentation, and CTGAN achieves a balance across all five criteria, making it suitable for general applications with moderate privacy concerns. As a result, our findings offer valuable insights for choosing the most suitable algorithm.

生成模型金融数据隐私保护数据合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。