arXiv:2608.12342cs.CLcs.LG2026-08ACL被引 2

测试大模型能否发现财报中的错误,发现它们在复杂场景下仍表现不佳。

Are Large Language Models Reliable Reviewers? A Benchmark for Error Detection in Financial Documents

论文配图:Are Large Language Models Reliable Reviewers? A Benchmark for Error Detection in Financial Documents
图 1 · 摘自论文原文
  • 构建首个金融错漏检测基准,覆盖9类真实场景
  • 超900份2025年未见数据,测试大模型在高阶推理下的表现
  • 提示微调可显著提升弱模型性能,适合金融合规研究者

确保财务文件的准确性对经济分析、监管合规和企业决策至关重要。尽管已有研究显示大语言模型(LLMs)在股价预测和财务分析等任务中表现良好,但其识别财务文档错误的能力尚未被充分探索。本文提出 extbf{FinED-Bench},首个面向金融错漏检测的公开基准,涵盖三个认知复杂度层级,覆盖九类真实金融场景,包含超过900份2025年发布的、现有语言模型未曾见过的文档。我们详述了基准构建过程,并评估了GPT-4o、Qwen3-14B等先进模型在此任务上的表现,该任务需结合金融领域知识与推理能力。实验结果表明,当前大模型在高复杂度情况下仍表现不佳;此外,监督微调能显著提升弱模型的表现。数据与代码已公开于https://github.com/hedyHe/FinED-Bench。

原文摘要 · Abstract (English)

Ensuring the accuracy of financial documents is critical for economic analysis, regulatory compliance, and corporate decision-making. Several studies have shown that Large Language Models (LLMs) perform well in many financial tasks, such as stock price movements and financial analytics. However, a critical task remains unexplored: the ability of LLMs to identify errors in financial documents. In this paper, we introduce \textbf{FinED-Bench}, the first publicly \textbf{Bench}mark for \textbf{Fin}ancial \textbf{E}rror \textbf{D}etection across three levels of cognitive complexity. FinED-Bench covers nine real-world financial scenarios, and includes over 900 documents reported in 2025 that are unseen by existing language models. We detail the benchmark construction process and evaluate several advanced LLMs (e.g., GPT-4o, Qwen3-14B) on this tasks, which requires both financial domain knowledge and reasoning capabilities. Experimental results show that current LLMs still struggle with this task, especially in high-complexity cases. Besides, supervised fine-tuning can significantly improve the performance of weaker LLMs on this task. Our data and code are available at https://github.com/hedyHe/FinED-Bench.

金融文本大模型评测错误检测财报分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。