arXiv:2507.06266q-fin.RMcs.AI2025-07被引 23

用机器学习提升企业审计效率,精准识别高风险与舞弊行为。

Machine Learning based Enterprise Financial Audit Framework and High Risk Identification

  • 采用随机森林等三种算法构建风险预测模型
  • 随机森林F1得分达0.9012,有效识别舞弊与合规异常
  • 适合关注智能审计与风险管理的企业及研究者

在全球经济不确定性背景下,财务审计对合规与风险防控愈发关键。传统人工审计因数据量大、业务结构复杂及欺诈手段演变而面临局限。本研究提出基于机器学习的企业财务审计框架与高风险识别方法,利用2020至2025年四大会计师事务所(埃森哲、普华永道、德勤、毕马威)的数据集,分析风险评估、合规违规与舞弊检测趋势。数据涵盖审计项目数、高风险案例、舞弊事件、合规违规、员工工作量及客户满意度等指标,反映审计行为与人工智能影响。评估支持向量机(SVM)、随机森林(RF)与K近邻(KNN)三种算法,其中随机森林在层级K折交叉验证中表现最优,F1-score达0.9012,显著提升舞弊与合规异常识别能力。特征重要性分析显示,审计频率、历史违规、员工工作量与客户评分是关键预测因子。研究建议采用随机森林为核心模型,通过特征工程优化,并实现风险实时监控。为现代企业智能审计与风险管理提供实践参考。

原文摘要 · Abstract (English)

In the face of global economic uncertainty, financial auditing has become essential for regulatory compliance and risk mitigation. Traditional manual auditing methods are increasingly limited by large data volumes, complex business structures, and evolving fraud tactics. This study proposes an AI-driven framework for enterprise financial audits and high-risk identification, leveraging machine learning to improve efficiency and accuracy. Using a dataset from the Big Four accounting firms (EY, PwC, Deloitte, KPMG) from 2020 to 2025, the research examines trends in risk assessment, compliance violations, and fraud detection. The dataset includes key indicators such as audit project counts, high-risk cases, fraud instances, compliance breaches, employee workload, and client satisfaction, capturing both audit behaviors and AI's impact on operations. To build a robust risk prediction model, three algorithms - Support Vector Machine (SVM), Random Forest (RF), and K-Nearest Neighbors (KNN) - are evaluated. SVM uses hyperplane optimization for complex classification, RF combines decision trees to manage high-dimensional, nonlinear data with resistance to overfitting, and KNN applies distance-based learning for flexible performance. Through hierarchical K-fold cross-validation and evaluation using F1-score, accuracy, and recall, Random Forest achieves the best performance, with an F1-score of 0.9012, excelling in identifying fraud and compliance anomalies. Feature importance analysis reveals audit frequency, past violations, employee workload, and client ratings as key predictors. The study recommends adopting Random Forest as a core model, enhancing features via engineering, and implementing real-time risk monitoring. This research contributes valuable insights into using machine learning for intelligent auditing and risk management in modern enterprises.

机器学习财务审计风险识别随机森林

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。