用无监督方法识别政府支出异常,提升审计效率与准确性。
Unsupervised Outlier Detection in Audit Analytics: A Case Study Using USA Spending Data
- 对比多种无监督算法检测联邦支出异常
- 混合策略比单一方法更准确,F1得分更高
- 适合关注财政监管与智能审计的研究者
本研究探讨无监督异常检测方法在审计分析中的有效性,以美国卫生与公共服务部(DHHS)的联邦支出数据为案例。采用基于直方图的异常评分(HBOS)、鲁棒主成分分析(Robust PCA)、最小协方差确定(MCD)和K近邻(KNN)等多种算法,识别财政支出中的异常模式。研究针对大规模政府数据中传统审计方法效率不足的问题,通过数据预处理、算法实现与性能评估(使用精确率、召回率、F1分数)进行验证。结果表明,结合多种检测策略的混合方法在复杂金融数据中显著提升了异常识别的鲁棒性与准确性。研究为审计分析领域提供了不同异常检测模型的比较见解,并展示了无监督学习在提升审计质量与效率方面的潜力。成果对审计人员、政策制定者及研究者具有重要参考价值。
原文摘要 · Abstract (English)
This study investigates the effectiveness of unsupervised outlier detection methods in audit analytics, utilizing USA spending data from the U.S. Department of Health and Human Services (DHHS) as a case example. We employ and compare multiple outlier detection algorithms, including Histogram-based Outlier Score (HBOS), Robust Principal Component Analysis (PCA), Minimum Covariance Determinant (MCD), and K-Nearest Neighbors (KNN) to identify anomalies in federal spending patterns. The research addresses the growing need for efficient and accurate anomaly detection in large-scale governmental datasets, where traditional auditing methods may fall short. Our methodology involves data preparation, algorithm implementation, and performance evaluation using precision, recall, and F1 scores. Results indicate that a hybrid approach, combining multiple detection strategies, enhances the robustness and accuracy of outlier identification in complex financial data. This study contributes to the field of audit analytics by providing insights into the comparative effectiveness of various outlier detection models and demonstrating the potential of unsupervised learning techniques in improving audit quality and efficiency. The findings have implications for auditors, policymakers, and researchers seeking to leverage advanced analytics in governmental financial oversight and risk management.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。