arXiv:2510.08948cs.IRcs.AI2025-10KDD被引 3

用知识库+大模型动态对抗电商欺诈,效率提升超3倍

SHERLOCK: Towards Dynamic Knowledge Adaptation in LLM-enhanced E-commerce Risk Management

  • 构建领域知识库,融合多源异构数据增强大模型推理
  • 实现82%专家认可率,日调查量提升386.7%
  • 自进化系统可自动修复性能衰退,适合风控场景持续迭代

有效的电商风险管控需深入分析案件以识别新兴欺诈模式,但人工调查依赖跨源异构数据关联,耗时耗力。尽管大语言模型(LLMs)在自动化分析中展现潜力,其应用受限于风险场景复杂性及长尾领域知识稀疏。为此,我们提出Sherlock框架,通过三大模块整合结构化领域知识与基于LLM的推理:首先,从异构知识源中提炼结构化专业知识构建领域知识库(KB);其次,设计双阶段检索增强生成策略,结合输入上下文增强与反射优化模块,充分挖掘KB以提升分析质量;最后,开发集成运营与标注平台,驱动自演化数据飞轮。通过实时知识库更新与周期性后训练逻辑对齐,实现系统持续演进以应对对抗漂移。京东线上A/B测试显示,Sherlock达成82%专家接受率(EAR),日调查吞吐量提升386.7%。90天评估表明,飞轮机制成功两次恢复因战术变化导致的性能衰减,通过自主模型更新将EAR上限提升约3.5%。

原文摘要 · Abstract (English)

Effective e-commerce risk management requires in-depth case investigations to identify emerging fraud patterns in highly adversarial environments. However, manual investigation typically requires analyzing the associations and couplings among multi-source heterogeneous data, a labor-intensive process that limits efficiency. While Large Language Models (LLMs) show promise in automating these analyses, their deployment is hindered by the complexity of risk scenarios and the sparsity of long-tail domain knowledge. To address these challenges, we propose Sherlock, a framework that integrates structured domain knowledge with LLM-based reasoning through three core modules. First, we construct a domain Knowledge Base (KB) by distilling structured expertise from heterogeneous knowledge sources. Second, we design a two-stage retrieval-augmented generation strategy tailored for case investigation, which combines input contextual augmentation with a Reflect & Refine module to fully leverage the KB for improved analysis quality. Finally, we develop an integrated platform for operations and annotation to drive a self-evolving data flywheel. By combining real-time hotfixes through KB updates with periodic logic alignment via post-training, we facilitate continuous system evolution to counteract adversarial drifts. Online A/B tests at JD dot com demonstrate that Sherlock achieves an 82% Expert Acceptance Rate (EAR) and a 386.7% increase in daily investigation throughput. An additional 90-day evaluation shows that the flywheel successfully recovers from performance decay caused by changing tactics twice, raising the EAR ceiling by around 3.5% through autonomous model updates.

大模型应用风险控制知识库自进化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。