用可追溯的推理提升虚假订单检测准确率,减少人工审核负担。
Traceable LLM Reasoning for Fake-Order Fraud Detection

- 通过语义统一模块将风险信号转为语言描述,让大模型理解。
- 在真实数据集上达到85.3%宏F1,优于基线2.7个百分点。
- 适合需要可解释性与自动化风控的平台,尤其关注效率与成本。
大规模在线到线下(O2O)服务平台的虚假订单欺诈检测仍面临严峻挑战,现有方法多依赖专家设计特征、决策黑箱且可解释性差。为此,我们提出DeepScrub,一种基于大语言模型(LLM)的强化学习框架,实现可追溯的欺诈检测推理。DeepScrub引入三项创新:第一,语义统一模块将异构风险信号转化为大模型可理解的文本描述;第二,基于风险控制语料库的持续预训练注入领域知识,并通过任务奖励联合评估预测准确性和推理质量;第三,提出SUggest-REflect(SURE)机制,融合专家反馈与模型自检,迭代优化推理路径。在真实世界虚假订单检测数据集上,DeepScrub取得85.3%的宏F1分数,优于最佳基线2.7个百分点。其任务优化的8B模型甚至超越32B模型,表明领域适配比模型规模更关键。在为期四周的线上试点中,该系统实现91.8%精确率与88.5%召回率,较初筛人工审核提升16.6和38.8个百分点,减少初筛人工工作量94%,年节省近百万人民币。结果表明,DeepScrub显著提升欺诈审查精度,降低人工负荷,并为生产环境风控流程提供可追溯证据。
原文摘要 · Abstract (English)
Detecting fake-order fraud at scale remains a critical challenge for large online-to-offline (O2O) service platforms, as existing approaches often rely on expert-designed features, produce black-box decisions, and provide limited interpretability. To address these limitations, we propose DeepScrub, a reinforcement learning framework built upon large language models (LLMs) for fake-order fraud detection with traceable reasoning. DeepScrub introduces three innovations. First, a semantic unification module converts heterogeneous risk signals into textual descriptions that LLMs can understand. Second, continued pre-training on risk-control corpora injects domain knowledge, and task rewards jointly evaluate prediction correctness and reasoning quality. Third, the SUggest-REflect (SURE) mechanism incorporates expert feedback and model self-checking to iteratively refine reasoning paths. On a real-world fake-order fraud detection dataset, DeepScrub achieves a macro-F1 score of 85.3%, outperforming the best baseline by 2.7 percentage points. Our task-optimized 8B model further surpasses a 32B model, showing that domain adaptation can matter more than model scale in this setting. In a four-week live pilot, DeepScrub achieved 91.8% precision and 88.5% recall, improving over first-stage human reviewers by 16.6 and 38.8 percentage points. It reduced first-stage manual review workload by 94% and saved nearly one million RMB annually. These results show that DeepScrub improves fraud review accuracy, reduces first-stage review workload, and provides traceable evidence for production risk-review workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。