arXiv:2507.16642cs.CLcs.AI2025-07中稿 · and published at B…

用大模型自动检查财报是否合规,开源模型表现亮眼。

Towards Automated Regulatory Compliance Verification in Financial Auditing with Large Language Models

  • 对比开源与商用大模型在财务合规验证中的表现。
  • Llama-2 70B 在识别不合规内容上优于所有商用模型。
  • GPT-4 在多语言场景下综合表现最佳,适合跨国审计。

财务审计历来是人力密集型工作,正面临转型。当前基于AI的解决方案能从财报中推荐符合会计准则的文本片段,但普遍缺乏对推荐内容是否真正符合法律要求的验证能力。本文评估公开可用的大语言模型(LLMs)在不同配置下对财务监管合规性的验证效率。重点对比Llama-2等开源模型与OpenAI GPT系列等专有模型的表现。实验基于德国普华永道(PwC Germany)提供的两个定制数据集。结果显示,开源的Llama-2 700亿参数模型在检测非合规或真负样本方面表现优异,超越所有专有模型;而如GPT-4等专有模型在多种场景下整体表现最优,尤其在非英语语境中优势明显。

原文摘要 · Abstract (English)

The auditing of financial documents, historically a labor-intensive process, stands on the precipice of transformation. AI-driven solutions have made inroads into streamlining this process by recommending pertinent text passages from financial reports to align with the legal requirements of accounting standards. However, a glaring limitation remains: these systems commonly fall short in verifying if the recommended excerpts indeed comply with the specific legal mandates. Hence, in this paper, we probe the efficiency of publicly available Large Language Models (LLMs) in the realm of regulatory compliance across different model configurations. We place particular emphasis on comparing cutting-edge open-source LLMs, such as Llama-2, with their proprietary counterparts like OpenAI's GPT models. This comparative analysis leverages two custom datasets provided by our partner PricewaterhouseCoopers (PwC) Germany. We find that the open-source Llama-2 70 billion model demonstrates outstanding performance in detecting non-compliance or true negative occurrences, beating all their proprietary counterparts. Nevertheless, proprietary models such as GPT-4 perform the best in a broad variety of scenarios, particularly in non-English contexts.

金融审计大模型合规验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。