arXiv:2512.02726cs.AI2025-12

用大模型检测会计分录中的舞弊,比传统方法更准且能解释。

AuditCopilot: Leveraging LLMs for Fraud Detection in Double-Entry Bookkeeping

  • 用大模型分析会计分录,自动识别异常模式。
  • 在真实和合成数据上,准确率显著高于规则和机器学习方法。
  • 生成自然语言解释,适合审计人员辅助决策。

审计师依赖凭证测试(JETs)检测税务账簿中的异常,但基于规则的方法会产生大量误报,难以发现细微异常。我们研究大语言模型(LLMs)能否作为双式记账中的异常检测工具。在合成数据和真实匿名账簿上,对LLaMA、Gemma等前沿大模型进行基准测试,对比传统JETs与机器学习基线。结果表明,大模型在各项指标上持续优于规则型JETs和经典机器学习方法,同时提供自然语言解释,增强可解释性。这些结果凸显了‘AI增强审计’的潜力,即人类审计员与基础模型协作,共同提升财务完整性。

原文摘要 · Abstract (English)

Auditors rely on Journal Entry Tests (JETs) to detect anomalies in tax-related ledger records, but rule-based methods generate overwhelming false positives and struggle with subtle irregularities. We investigate whether large language models (LLMs) can serve as anomaly detectors in double-entry bookkeeping. Benchmarking SoTA LLMs such as LLaMA and Gemma on both synthetic and real-world anonymized ledgers, we compare them against JETs and machine learning baselines. Our results show that LLMs consistently outperform traditional rule-based JETs and classical ML baselines, while also providing natural-language explanations that enhance interpretability. These results highlight the potential of \textbf{AI-augmented auditing}, where human auditors collaborate with foundation models to strengthen financial integrity.

财务审计大模型应用异常检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。