提出动态数据模拟框架,真实评估迁移学习在金融欺诈检测中的表现。
Evaluating Transfer Learning Methods on Real-World Data Streams: A Case Study in Financial Fraud Detection
- 构建随时间变化的数据可用性模拟框架,支持动态评估。
- 通过重采样与领域变换生成多样真实场景,覆盖概念漂移等挑战。
- 适用于金融、支付等数据流持续变化的现实场景模型部署决策。
当目标领域数据有限时,迁移学习(TL)可通过相关数据丰富的领域预训练模型。然而,现有方法多基于静态的数据量假设,与实际中数据及标签随时间波动的情况不符。为此,本文提出一种数据操作框架,可(1)模拟随时间变化的数据可用性,(2)通过重采样生成多个领域,(3)引入真实域间差异,如时间相关的协变量和概念漂移。该框架能生成大量真实场景变体,揭示算法在动态环境中的潜在行为。我们以某公司真实信用卡支付数据集为案例研究,并在公开的银行账户欺诈(BAF)数据集上验证其有效性。该框架为迁移学习方法在时间演化数据中的评估提供了新范式,有助于提升实际部署中的模型决策质量。
原文摘要 · Abstract (English)
When the available data for a target domain is limited, transfer learning (TL) methods can be used to develop models on related data-rich domains, before deploying them on the target domain. However, these TL methods are typically designed with specific, static assumptions on the amount of available labeled and unlabeled target data. This is in contrast with many real world applications, where the availability of data and corresponding labels varies over time. Since the evaluation of the TL methods is typically also performed under the same static data availability assumptions, this would lead to unrealistic expectations concerning their performance in real world settings. To support a more realistic evaluation and comparison of TL algorithms and models, we propose a data manipulation framework that (1) simulates varying data availability scenarios over time, (2) creates multiple domains through resampling of a given dataset and (3) introduces inter-domain variability by applying realistic domain transformations, e.g., creating a variety of potentially time-dependent covariate and concept shifts. These capabilities enable simulation of a large number of realistic variants of the experiments, in turn providing more information about the potential behavior of algorithms when deployed in dynamic settings. We demonstrate the usefulness of the proposed framework by performing a case study on a proprietary real-world suite of card payment datasets. Given the confidential nature of the case study, we also illustrate the use of the framework on the publicly available Bank Account Fraud (BAF) dataset. By providing a methodology for evaluating TL methods over time and in realistic data availability scenarios, our framework facilitates understanding of the behavior of models and algorithms. This leads to better decision making when deploying models for new domains in real-world environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。