arXiv:2503.16974q-fin.GNcs.AI2025-03被引 40

首次评估大模型在金融会计任务中的输出一致性,发现复杂任务差异大但聚合可显著提升

Assessing Consistency and Reproducibility in the Outputs of Large Language Models: Evidence Across Diverse Finance and Accounting Tasks

  • 通过50次独立运行测试5类金融任务,对比GPT-3.5、GPT-4o-mini和GPT-4o三模型表现
  • 二分类与情感分析几乎完全可复现,复杂任务如生成和预测变异性更高,但聚合3-5次后一致性大幅提升
  • 新模型聚合后准确率提高,且统计推断对不一致不敏感,适合金融研究使用

本研究首次全面评估大语言模型(LLM)在金融与会计研究中输出的一致性与可复现性。通过在五类常见任务(分类、情感分析、摘要、文本生成、预测)上进行50次独立实验,使用三个OpenAI模型(GPT-3.5-turbo、GPT-4o-mini、GPT-4o)生成超过340万条输出,覆盖包括MD&A、FOMC声明、财经新闻、财报电话会议记录和财务报表在内的多样化金融文本。结果表明,一致性存在显著任务依赖性:二分类与情感分析近乎完全可复现,而复杂任务变异性更高。更先进模型并未始终表现出更好一致性,呈现任务特异性模式。大模型在一致性上显著优于人类专家,且在人类分歧较大时仍保持高一致度。简单聚合3-5次运行结果可大幅改善一致性;对新版模型,聚合还能提升情感分析准确率。模拟分析显示,尽管输出存在可测量不一致,下游统计推断依然高度稳健。这些发现回应了所谓“G-hacking”(即从多次生成中选择有利结果)的担忧,表明其在金融会计任务中风险相对较低。

原文摘要 · Abstract (English)

This study provides the first comprehensive assessment of consistency and reproducibility in Large Language Model (LLM) outputs in finance and accounting research. We evaluate how consistently LLMs produce outputs given identical inputs through extensive experimentation with 50 independent runs across five common tasks: classification, sentiment analysis, summarization, text generation, and prediction. Using three OpenAI models (GPT-3.5-turbo, GPT-4o-mini, and GPT-4o), we generate over 3.4 million outputs from diverse financial source texts and data, covering MD&As, FOMC statements, finance news articles, earnings call transcripts, and financial statements. Our findings reveal substantial but task-dependent consistency, with binary classification and sentiment analysis achieving near-perfect reproducibility, while complex tasks show greater variability. More advanced models do not consistently demonstrate better consistency and reproducibility, with task-specific patterns emerging. LLMs significantly outperform expert human annotators in consistency and maintain high agreement even where human experts significantly disagree. We further find that simple aggregation strategies across 3-5 runs dramatically improve consistency. We also find that aggregation may come with an additional benefit of improved accuracy for sentiment analysis when using newer models. Simulation analysis reveals that despite measurable inconsistency in LLM outputs, downstream statistical inferences remain remarkably robust. These findings address concerns about what we term "G-hacking," the selective reporting of favorable outcomes from multiple generative AI runs, by demonstrating that such risks are relatively low for finance and accounting tasks.

大模型评估金融会计一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。