arXiv:2607.09641cs.LGcs.AI2026-07

用多目标强化学习破解金融反欺诈中的误判与漏判难题

Semantic Pareto-DQN: A Multi-Objective Reinforcement Learning Framework for Financial Anomaly Detection

  • 将交易特征转为自然语言描述,生成不变尺度的状态表示
  • 同时优化检测效果、操作摩擦和语义发现,召回率显著提升
  • 适合需要平衡风控与用户体验的金融场景

金融异常检测面临极端类别不平衡问题,传统单目标算法易出现‘欺诈坍塌’,过度偏向多数类,难以兼顾异常拦截与客户体验。为避免数据重采样的失真,我们提出语义帕累托DQN(Semantic Pareto-DQN),一种多目标强化学习框架。该方法通过大语言模型将异构交易特征融合为连贯的自然语言叙述,生成鲁棒且尺度不变的状态表示。智能体优化一个向量奖励,显式解耦金融有效性、运营摩擦与语义发现。通过映射连续帕累托前沿,系统可动态权衡漏报异常与误报的不对称成本。在电商欺诈与UCI信用卡数据集上的实证评估表明,语义帕累托DQN成功打破零召回陷阱,相比标量基线,在少数类召回率上表现更优,为在控制运营摩擦的前提下实现异常发现提供了新路径。

原文摘要 · Abstract (English)

Financial anomaly detection suffers from extreme class imbalance, causing traditional single-objective algorithms to exhibit ``fraud collapse'', defaulting to the majority class and failing to balance anomaly interdiction with customer friction. To overcome this without distortive data resampling, we propose the Semantic Pareto-DQN, a multi-objective reinforcement learning framework. Our approach synthesizes heterogeneous transaction features into cohesive natural-language narratives, encoded by large language models, thereby producing a robust, scale-invariant state representation. The agent optimizes a vectorial reward that explicitly decouples financial efficacy, operational friction, and semantic discovery. By mapping the continuous Pareto frontier, the system dynamically navigates the asymmetric costs of missed anomalies versus false positives. Empirical evaluations across E-Commerce fraud and UCI Credit datasets show that semantic Pareto-DQN successfully shatters the zero-recall trap. It achieves superior minority-class recall compared to scalarized baselines, providing an alternative to trade bounded operational friction for financial anomaly discovery.

反欺诈强化学习多目标金融风控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。