用多智能体融合财务证据,显式建模不确定性和冲突,提升审计风险预测准确率。
Multi-Agent Framework for Audit Risk Assessment with Explicit Uncertainty and Evidence Conflict Modeling
- 三类智能体分别分析文本、财务比率和异常信号,输出带置信度的风险评分
- 在3200个公司年数据上,AUROC达0.782,PR-AUC为0.341,优于多种基线方法
- 能识别证据冲突模式,提供可解释的审计风险信号,适合审计人员使用
审计风险评估日益依赖融合异构证据来源,但现有方法通常仅输出点预测,无法量化不同证据流的一致性。本文提出UMAR(不确定性感知多智能体风险评估)框架,包含三类专用智能体:MD&A文本智能体、财务比率智能体和异常检测智能体(CAM),各自独立生成带校准不确定性估计的风险评分。基于Dempster-Shafer证据理论的不确定性聚合器融合这些评分,并显式度量智能体间冲突。在涵盖3,200个公司年观测值的美国样本(2019–2023年SEC 10-K文件)上评估,以财务重述为目标标签。实验结果表明,UMAR达到AUROC 0.782、PR-AUC 0.341,优于逻辑回归、XGBoost、FinBERT以及单/双智能体大模型基线。其预期校准误差最低(ECE = 0.052),并识别出与实际重述风险相关的证据冲突模式,为审计师提供潜在可操作且可解释的风险信号。
原文摘要 · Abstract (English)
Audit risk assessment increasingly benefits from combining heterogeneous evidence sources, yet existing approaches typically produce point predictions without quantifying how well different evidence streams agree. We propose UMAR (Uncertainty-Aware Multi-Agent Risk Assessment), a framework that employs three specialized agents: an MD&A Text Agent, a Financial Ratio Agent, and a CAM Agent, each producing independent risk scores with calibrated uncertainty estimates. An Uncertainty Aggregator based on Dempster-Shafer evidence theory fuses these scores while explicitly measuring inter-agent conflict. We evaluate UMAR on a U.S. dataset of 3,200 firm-year observations from SEC 10-K filings (2019-2023), with financial restatement as the target label. Experimental results show that UMAR achieves an AUROC of 0.782 and a PR-AUC of 0.341, outperforming logistic regression, XGBoost, FinBERT, and single-agent and dual-agent LLM baselines. UMAR attains the lowest expected calibration error (ECE = 0.052) among all methods and identifies evidence-conflict patterns that correlate with actual restatement risk, offering auditors potentially actionable and interpretable risk signals.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。