用大模型自动优化强化学习奖励函数,提升多阶段支付欺诈检测效果
LLM-Enhanced Self-Evolving Reinforcement Learning for Multi-Step E-Commerce Payment Fraud Risk Detection
- 用大模型迭代优化强化学习的奖励函数设计
- 在真实数据上实现更高欺诈检测准确率和零样本能力
- 适合工业级强化学习应用开发者参考
本文提出一种融合强化学习(RL)与大语言模型(LLMs)的新方法,用于电商支付欺诈检测。将交易风险建模为多步马尔可夫决策过程(MDP),利用强化学习优化跨多个支付阶段的风险识别。传统奖励函数设计需大量人工经验,而大语言模型凭借其推理与编码能力,可有效改进该过程。本方法通过大模型迭代优化奖励函数,在真实世界数据上验证了其有效性、鲁棒性与长期稳定性,展现出优于传统方法的欺诈检测性能,并具备零样本能力,彰显大语言模型在工业强化学习中的潜力。
原文摘要 · Abstract (English)
This paper presents a novel approach to e-commerce payment fraud detection by integrating reinforcement learning (RL) with Large Language Models (LLMs). By framing transaction risk as a multi-step Markov Decision Process (MDP), RL optimizes risk detection across multiple payment stages. Crafting effective reward functions, essential for RL model success, typically requires significant human expertise due to the complexity and variability in design. LLMs, with their advanced reasoning and coding capabilities, are well-suited to refine these functions, offering improvements over traditional methods. Our approach leverages LLMs to iteratively enhance reward functions, achieving better fraud detection accuracy and demonstrating zero-shot capability. Experiments with real-world data confirm the effectiveness, robustness, and resilience of our LLM-enhanced RL framework through long-term evaluations, underscoring the potential of LLMs in advancing industrial RL applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。