将可解释AI融入大数据反欺诈系统,提升透明度与可信度。
Explainable AI in Big Data Fraud Detection
- 结合分布式存储、流处理与多种检测模型构建分析框架
- 评估LIME、SHAP等XAI方法在大规模场景下的效果与局限
- 提出融合实时反馈的可扩展解释机制,适合风控与合规场景
大数据在金融、保险和网络安全中日益关键,支撑大规模风险评估与欺诈检测。但自动化分析带来的透明度不足、合规风险与信任缺失问题亟待解决。本文探讨可解释人工智能(XAI)如何集成到大数据反欺诈分析流程中。综述了大数据的核心特征及主流分析工具,包括分布式存储、流式处理平台以及异常检测、图模型和集成分类器等欺诈检测模型。系统回顾了LIME、SHAP、反事实解释和注意力机制等广泛使用的XAI方法,分析其在规模化部署中的优缺点。基于研究发现,识别出可扩展性、实时处理以及图结构与时间序列模型解释能力方面的关键空白。为此,提出一个概念框架,将可扩展的大数据基础设施与上下文感知的解释机制及人工反馈相结合。最后,展望可扩展XAI、隐私保护解释与标准化评估方法等开放研究方向。
原文摘要 · Abstract (English)
Big Data has become central to modern applications in finance, insurance, and cybersecurity, enabling machine learning systems to perform large-scale risk assessments and fraud detection. However, the increasing dependence on automated analytics introduces important concerns about transparency, regulatory compliance, and trust. This paper examines how explainable artificial intelligence (XAI) can be integrated into Big Data analytics pipelines for fraud detection and risk management. We review key Big Data characteristics and survey major analytical tools, including distributed storage systems, streaming platforms, and advanced fraud detection models such as anomaly detectors, graph-based approaches, and ensemble classifiers. We also present a structured review of widely used XAI methods, including LIME, SHAP, counterfactual explanations, and attention mechanisms, and analyze their strengths and limitations when deployed at scale. Based on these findings, we identify key research gaps related to scalability, real-time processing, and explainability for graph and temporal models. To address these challenges, we outline a conceptual framework that integrates scalable Big Data infrastructure with context-aware explanation mechanisms and human feedback. The paper concludes with open research directions in scalable XAI, privacy-aware explanations, and standardized evaluation methods for explainable fraud detection systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。