arXiv:2601.22949cs.CLcs.CR2026-01被引 1

让大模型自动推理诈骗图谱,提升检测精度与速度。

Autonomous Chain-of-Thought Distillation for Graph-Based Fraud Detection

  • 大模型自主生成多跳推理路径,增强图文结构理解。
  • 比顶尖方法高8.8%的AUPRC,训练速度提升1066倍。
  • 适合需要高效精准反欺诈的工业场景使用。

基于文本属性图(TAGs)的欺诈检测需联合建模丰富的文本语义与关系依赖。现有基于大模型的图神经网络方法受限于预设提示词和分离的训练流程,削弱了推理自主性与语义-结构对齐。本文提出FraudCoT框架,通过自主、图感知的链式思维(CoT)推理与可扩展的LLM-GNN联合训练,实现统一建模。为突破固定提示限制,引入欺诈感知的有选择性CoT蒸馏机制,生成多样化推理路径并融入节点文本,为GNN提供多跳语义与结构线索。同时设计高效非对称联合训练策略,实现端到端优化,显著降低计算开销。在公开与工业基准上实验表明,FraudCoT相较现有方法最高提升8.8% AUPRC,训练吞吐量提升达1,066倍,大幅改善检测性能与效率。

原文摘要 · Abstract (English)

Graph-based fraud detection on text-attributed graphs (TAGs) requires jointly modeling rich textual semantics and relational dependencies. However, existing LLM-enhanced GNN approaches are constrained by predefined prompting and decoupled training pipelines, limiting reasoning autonomy and weakening semantic-structural alignment. We propose FraudCoT, a unified framework that advances TAG-based fraud detection through autonomous, graph-aware chain-of-thought (CoT) reasoning and scalable LLM-GNN co-training. To address the limitations of predefined prompts, we introduce a fraud-aware selective CoT distillation mechanism that generates diverse reasoning paths and enhances semantic-structural understanding. These distilled CoTs are integrated into node texts, providing GNNs with enriched, multi-hop semantic and structural cues for fraud detection. Furthermore, we develop an efficient asymmetric co-training strategy that enables end-to-end optimization while significantly reducing the computational cost of naive joint training. Extensive experiments on public and industrial benchmarks demonstrate that FraudCoT achieves up to 8.8% AUPRC improvement over state-of-the-art methods and delivers up to 1,066x speedup in training throughput, substantially advancing both detection performance and efficiency.

反欺诈图神经网络大模型推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。