构建首个金融反欺诈异构图基准,真实模拟多实体复杂关系。
FinFraudBench: A Heterogeneous Graph Benchmark for Financial Fraud Detection

- 用六类实体、十四种边类型建模真实金融网络关系。
- 数据集达890万条边,含极端不平衡标签和自然欺诈率。
- 提供标准化评估体系,助力实际场景模型测试。
数字金融系统日益复杂,使欺诈检测从孤立交易判断转向基于关联实体的风险推理。现有图模型虽有进展,但公开基准在两方面与真实金融系统脱节:一是多简化为同质或单一节点类型的多关系图,忽略金融数据的多元主体与多关系本质;二是缺乏大规模、带真实运营条件(如极端类别不平衡、标签稀缺)的异构图数据集,难以评估方法实际效果。为此,我们提出FinFraudBench,一个面向金融反欺诈的异构图基准。包含两个异构图数据集(CreditCard-Fraud与BankTrans-Fraud),最大规模达8923万条有向类型边、899万节点,涵盖六类金融实体、十四种有向边类型,且自然欺诈率符合部署约束。我们建立了覆盖排序与不平衡敏感分类指标的标准评估协议,并对代表性基线进行评测。大量实验揭示了当前方法的局限,指明未来研究方向。数据集已开源:https://anonymous.4open.science/r/FinFraudBench-B002。
原文摘要 · Abstract (English)
The increasing complexity of digital financial systems has reshaped financial fraud detection from isolated transaction classification into relational risk reasoning over interconnected financial entities. This shift has motivated graph-based fraud detection, where models identify fraudulent nodes by exploiting dependencies among customers, cards, merchants, categories, and locations. However, despite rapid progress in graph-based methods, existing public benchmarks remain misaligned with real-world financial systems in two important aspects. First, they often simplify financial ecosystems into homogeneous or single-node-type multi-relational graphs, failing to preserve the multi-entity and multi-relational nature of financial data. Second, they rarely provide large-scale heterogeneous financial graph datasets with realistic operating conditions such as extreme class imbalance and limited label availability, making it difficult to assess the practical effectiveness of current methods. To address these gaps, we present FinFraudBench, a heterogeneous graph benchmark for financial fraud detection. FinFraudBench contains two heterogeneous graph datasets (CreditCard-Fraud and BankTrans-Fraud) with up to 8.99M nodes and 89.23M directed typed edges. Each dataset preserves six financial entity types, fourteen directed edge types, and natural fraud rates that mirror deployment constraints. With these datasets, we establish a standardized evaluation protocol covering both ranking and imbalance-sensitive classification metrics, and evaluate representative baselines. Extensive experiments yield empirical insights into current methods' limitations and suggest promising avenues for future research. FinFraudBench is available at https://anonymous.4open.science/r/FinFraudBench-B002.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。