无需共享模型或数据,单轮通信实现跨机构财务异常检测
Anomaly Detection in Double-entry Bookkeeping Data by Federated Learning System with Non-model Sharing Approach
- 通过数据协同分析将原始凭证转为安全中间表示
- 在8家医疗机构数据上优于本地训练和传统联邦学习方法
- 适合医疗、审计等敏感数据场景下的智能风控系统
异常检测在财务审计中至关重要,但需要来自多个机构的大规模数据。然而,分录数据高度敏感,难以直接共享。现有基于模型共享的联邦学习方法需多轮通信且要求设备连接外部网络,不符合敏感数据安全规范。为此,本文提出一种非模型共享型联邦学习框架,通过降维将原始分录数据转化为安全中间表示,并构建协作表示用于训练异常检测自编码器。该方法无需暴露原始数据,也无需设备接入外部网络,仅需单轮通信。在8家医疗机构的真实与合成数据上评估显示,该框架不仅优于单一机构本地训练基线,还在非独立同分布(non-i.i.d.)环境下超越FedAvg和FedProx等模型共享联邦学习方法,有效平衡了知识融合与数据隐私保护,推动智能化审计系统落地。
原文摘要 · Abstract (English)
Anomaly detection is crucial in financial auditing, and effective detection requires large volumes of data from multiple organizations. However, journal entry data is highly sensitive, making it infeasible to share them directly across audit firms. To address this challenge, journal entry anomaly detection methods based on model share-type federated learning (FL) have been proposed. These methods require multiple rounds of communication with external servers to exchange model parameters, which necessitates connecting devices storing confidential data to external networks -- a practice not recommended for sensitive data such as journal entries. To overcome these limitations, a novel anomaly detection framework based on data collaboration (DC) analysis, a non-model share-type FL approach, is proposed. The method first transforms raw journal entry data into secure intermediate representations via dimensionality reduction and then constructs a collaboration representation used to train an anomaly detection autoencoder. Notably, the approach does not require raw data to be exposed or devices to be connected to external networks, and the entire process needs only a single round of communication. The proposed method was evaluated on both synthetic and real-world journal entry data collected from eight healthcare organizations. The experimental results demonstrated that the framework not only outperforms the baseline trained on individual data but also achieves higher detection performance than model-sharing FL methods such as FedAvg and FedProx, particularly under non-i.i.d. settings that simulate practical audit environments. This study addresses the critical need to integrate organizational knowledge while preserving data confidentiality, contributing to the development of practical intelligent auditing systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。