arXiv:2601.06707cs.CL2026-01

评测大模型在会计推理任务中的表现,发现提示工程影响显著。

Evaluating Accounting Reasoning Capabilities of Large Language Models

  • 构建会计推理评估框架,基于主流大模型训练数据特征
  • GPT-4表现最佳,但整体仍难满足企业实际需求
  • 适合关注大模型在专业领域应用的研究者与开发者

大型语言模型正在改变多个领域的学习、认知与研究方式。将它们有效融入会计等专业领域,是企业数字化转型的关键挑战。为此,我们定义了垂直领域的会计推理,并基于代表性GLM模型的训练数据特征,提出评估标准,支持对会计推理能力的系统性研究,并提供性能改进基准。利用该框架,我们评估了GLM-6B、GLM-130B、GLM-4和OpenAI GPT-4在会计推理任务上的表现。结果表明,提示设计显著影响模型性能,其中GPT-4展现出最强能力。尽管取得进展,当前模型仍不足以应对真实企业会计场景,亟需进一步优化以释放其实际应用潜力。

原文摘要 · Abstract (English)

Large language models are transforming learning, cognition, and research across many fields. Effectively integrating them into professional domains, such as accounting, is a key challenge for enterprise digital transformation. To address this, we define vertical domain accounting reasoning and propose evaluation criteria derived from an analysis of the training data characteristics of representative GLM models. These criteria support systematic study of accounting reasoning and provide benchmarks for performance improvement. Using this framework, we evaluate GLM-6B, GLM-130B, GLM-4, and OpenAI GPT-4 on accounting reasoning tasks. Results show that prompt design significantly affects performance, with GPT-4 demonstrating the strongest capability. Despite these gains, current models remain insufficient for real-world enterprise accounting, indicating the need for further optimization to unlock their full practical value.

大模型评估会计推理提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。