专为审计设计的开源大模型,提升审计效率与准确性。
AuditWen:An Open-Source Large Language Model for Audit
- 基于Qwen微调,构建28k条审计任务指令数据集。
- 在3000条评测指令上表现优于其他大模型。
- 适合审计人员快速处理文档与问答任务。
智能审计是人工智能时代审计实践的重要进展,显著提升审计质量和效率。尽管大语言模型(LLM)潜力巨大,但通用模型在审计领域存在专业知识不足和数据偏见问题。为此,本文提出AuditWen,一个通过在15个审计任务、3个层级上构建28,000条指令数据并微调Qwen而得到的开源审计专用大模型。我们首先梳理了审计中LLM的应用场景并提炼开发需求,随后构建涵盖关键审计任务的3,000条评测基准。实验表明,AuditWen在信息抽取、问答理解和文档生成任务中均表现优异,具备直接应用于审计工作的价值。
原文摘要 · Abstract (English)
Intelligent auditing represents a crucial advancement in modern audit practices, enhancing both the quality and efficiency of audits within the realm of artificial intelligence. With the rise of large language model (LLM), there is enormous potential for intelligent models to contribute to audit domain. However, general LLMs applied in audit domain face the challenges of lacking specialized knowledge and the presence of data biases. To overcome these challenges, this study introduces AuditWen, an open-source audit LLM by fine-tuning Qwen with constructing instruction data from audit domain. We first outline the application scenarios for LLMs in the audit and extract requirements that shape the development of LLMs tailored for audit purposes. We then propose an audit LLM, called AuditWen, by fine-tuning Qwen with constructing 28k instruction dataset from 15 audit tasks and 3 layers. In evaluation stage, we proposed a benchmark with 3k instructions that covers a set of critical audit tasks derived from the application scenarios. With the benchmark, we compare AuditWen with other existing LLMs from information extraction, question answering and document generation. The experimental results demonstrate superior performance of AuditWen both in question understanding and answer generation, making it an immediately valuable tool for audit.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。