用百亿级图学习提升微信支付信用风险识别能力
Empowering Credit Risk Detection in Weixin Pay with Billion-Scale Deep Graph Learning

- 基于重叠子图构建风险感知框架,兼顾负载均衡与拓扑完整性
- 预算约束采样保留长尾风险节点,有效捕获关键欺诈传播路径
- 跨子图表示对齐机制解决局部特征不一致问题,适合工业级风控场景
信用风险检测,特别是防范个体欺诈,对维护数字金融生态稳定至关重要。在数十亿用户中精准识别信用欺诈,有助于降低金融损失并保障普惠金融服务的可持续性。由于信用欺诈风险常隐藏于异构的用户-风险图中,图神经网络(GNN)通过捕捉复杂依赖关系成为有效的风险挖掘工具。为突破工业级GNN的可扩展性瓶颈,基于子图的分布式训练不可或缺。然而,现有策略常为负载均衡牺牲拓扑完整性,可能切断对风险传播至关重要的长尾证据链。重叠子图虽可恢复被切断的风险上下文,但引入冗余与噪声,且忽略不同局部子图间的表征对齐。本文提出一种面向大规模信用风险检测的风险感知重叠子图学习框架。首先构建基础划分以保证负载均衡;随后采用预算约束采样,选择具信息量的长尾节点,保留关键风险扩散模式的同时过滤噪声;为缓解表征不一致,设计跨子图一致性对齐机制,通过对重叠节点施加对齐约束,将局部表示统一至全局一致的潜在空间。在微信支付生产数据集上的大量实验表明,本模型显著优于现有策略,为工业级图学习提供了可扩展且高效的解决方案。
原文摘要 · Abstract (English)
Credit risk detection, particularly mitigating individual fraud, is crucial for maintaining the stability of digital financial ecosystems. Accurately identifying credit fraud among billions of users is critical for minimizing financial losses and safeguarding the sustainability of inclusive financial services. Given that credit fraud risks are often concealed within heterogeneous user-risk graphs, Graph Neural Networks (GNNs) have emerged as an effective tool for risk mining by capturing complex dependencies. To address the scalability bottleneck of industrial GNNs, distributed training based on subgraphs is indispensable. However, existing strategies often compromise topological integrity for load balancing. This can be catastrophic for risk detection, as it indiscriminately severs the long-tail evidence chains essential for risk propagation. Overlapping subgraphs can restore severed risk contexts but inevitably introduce redundancy and noise, while overlooking the representation alignment across different local subgraphs. In this paper, we propose a risk-aware overlapping subgraph learning framework for large-scale credit risk detection. We first construct base partitions to ensure load balance. Then, we perform budget-constrained sampling that selects informative long-tail nodes, thereby preserving critical risk diffusion patterns while filtering out noise. To mitigate representation inconsistency, we design a cross-subgraph consistency alignment mechanism. By enforcing alignment constraints on the overlapping nodes, we harmonize the local representations into a globally consistent latent space. Extensive experiments on Weixin Pay's production dataset demonstrate that our model significantly outperforms existing strategies for risk detection, offering a scalable and effective solution for industrial graph learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。