arXiv:2512.22443cs.CL2025-12

提出会计推理新概念,评估大模型在会计任务中的表现。

Accounting Reasoning in Large Language Models: Concepts, Evaluation, and Empirical Analysis

  • 定义会计推理并构建评估框架,基于GLM系列模型训练数据设计标准。
  • 测试GLM-6B、GLM-130B、GLM-4和GPT-4,GPT-4表现最佳但仍不足。
  • 适合关注大模型在专业领域应用的开发者与研究者参考。

大型语言模型(LLMs)正重塑多个领域的学习范式、认知过程与研究方法。随着其应用扩展,如何有效融入专业领域并明确其在特定场景中的角色,已成为企业数字化转型与社会发展的关键挑战。在会计领域,成功整合需系统理解模型的领域专用推理能力。本文提出会计推理概念,基于典型GLM系列模型的训练数据特征,构建评估标准,为研究会计导向推理范式提供基础,并设立性能评估与改进基准。在此框架下,我们评估了GLM-6B、GLM-130B、GLM-4和GPT-4等代表性模型在多种会计推理任务上的表现。实验显示,提示工程可不同程度提升各模型性能,其中GPT-4展现出最强的整体会计推理能力。然而结果表明,当前模型仍不足以支撑真实会计应用场景,尤其在企业级部署中仍需进一步优化,以充分释放其在该领域的潜力。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly reshaping learning paradigms, cognitive processes, and research methodologies across diverse domains. As their adoption expands, effectively integrating LLMs into professional fields and clarifying their role in domain-specific applications has become a key challenge for enterprise digital transformation and broader societal development. In the accounting domain, successful integration requires a systematic understanding of LLMs' domain-specific reasoning capabilities. In this study, we introduce the concept of accounting reasoning and propose a set of evaluation criteria grounded in an analysis of the training data characteristics of representative GLM-series models. These criteria establish a foundation for studying accounting-oriented reasoning paradigms and provide benchmarks for assessing and improving model performance. Building on this framework, we evaluate several representative LLMs, including GLM-6B, GLM-130B, GLM-4, and GPT-4, across a range of accounting reasoning tasks. Our experimental results show that prompt engineering strategies can yield varying degrees of performance improvement across models, with GPT-4 demonstrating the strongest overall accounting reasoning capability. Nevertheless, the results indicate that current LLMs remain insufficient for real-world accounting applications. In particular, further optimization is required for deployment in enterprise-level accounting scenarios to fully realize the potential value of LLMs in this domain.

会计推理大模型评估提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。