现有合成数据生成器无法保留真实欺诈行为模式,导致风控系统失效。
Synthetic Tabular Generators Fail to Preserve Behavioral Fraud Patterns: A Benchmark on Temporal, Velocity, and Multi-Account Signals
- 提出行为保真度评估维度,检测生成数据是否保留时间、序列与多账户行为特征
- 四种欺诈模式中,4种生成器在真实数据上表现差24.4至99.7倍
- 适用于金融、医疗等需分析实体行为的领域,开源评估框架
我们提出行为保真度——一种合成表格数据的新评估维度,用于衡量生成数据是否保留了真实世界实体活动中的时间、序列与结构行为模式。现有框架仅评估统计保真度(边际分布与相关性)和下游效用(分类器AUROC),但未测试实际风控系统依赖的行为信号。我们构建了涵盖事件间隔、爆发结构、多账户图模式和速度规则触发率的四类欺诈行为模式(P1-P4);定义了校准至真实数据噪声基线的退化比率(1.0=匹配真实变异,k=k倍更差);证明行独立生成器(主流范式)在结构上无法复现P3图模式(命题1),且产生非正的单实体事件间隔自相关(命题2),使欺诈序列的正向爆发指纹不可实现。我们在IEEE-CIS Fraud Detection与Amazon Fraud Dataset上评测CTGAN、TVAE、GaussianCopula和TabularARGN,全部表现严重不足:在IEEE-CIS上退化比率24.4倍(TVAE)至39.0倍(GaussianCopula);在Amazon FDB上,行独立生成器达81.6-99.7倍,而TabularARGN为17.2倍。我们记录了各生成器的失败模式及修复方案。P1-P4框架可推广至医疗、网络安全等实体级序列表格数据场景。评估框架已开源。
原文摘要 · Abstract (English)
We introduce behavioral fidelity -- a third evaluation dimension for synthetic tabular data that measures whether generated data preserves the temporal, sequential, and structural behavioral patterns that distinguish real-world entity activity. Existing frameworks evaluate statistical fidelity (marginal distributions and correlations) and downstream utility (classifier AUROC on synthetic-trained models), but neither tests for the behavioral signals that operational detection and analysis systems actually rely on. We formalize a taxonomy of four behavioral fraud patterns (P1-P4) covering inter-event timing, burst structure, multi-account graph motifs, and velocity-rule trigger rates; define a degradation ratio metric calibrated to a real-data noise floor (1.0 = matches real variability, k = k-times worse); and prove that row-independent generators -- the dominant paradigm -- are structurally incapable of reproducing P3 graph motifs (Proposition 1) and produce non-positive within-entity IET autocorrelation (Proposition 2), making the positive burst fingerprint of fraud sequences unachievable regardless of architecture or training data size. We benchmark CTGAN, TVAE, GaussianCopula, and TabularARGN on IEEE-CIS Fraud Detection and the Amazon Fraud Dataset. All four fail severely: on IEEE-CIS composite degradation ratios range from 24.4x (TVAE) to 39.0x (GaussianCopula); on Amazon FDB, row-independent generators score 81.6-99.7x, while TabularARGN achieves 17.2x. We document generator-specific failure modes and their resolutions. The P1-P4 framework extends to any domain with entity-level sequential tabular data, including healthcare and network security. We release our evaluation framework as open source.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。