arXiv:2603.07316cs.AI2026-03被引 3

测试大模型处理金融表格的推理能力,发现其在复杂数据上表现不佳。

FinSheet-Bench: From Simple Lookups to Complex Reasoning, Where LLMs Break on Financial Spreadsheets

  • 构建合成金融数据集,评估大模型对表格文本与数值推理的能力
  • 最佳模型准确率仅82.4%,复杂表格上降至48.6%
  • 提示架构需分离理解与计算,才能提升金融场景可靠性

尽管大语言模型(LLMs)能加速另类投资尽职调查中的文本任务,但在从复杂财务表格中准确提取和推理结构化数据方面仍存在差距。由于私募股权数据室保密,缺乏真实行业基金组合数据用于基准测试。为此,我们提出了FinSheet-Bench,一个基于真实私募股权基金结构构建的合成财务组合数据集,旨在评估大模型在文本序列化表格问答和数值推理任务上的表现。我们对OpenAI、Google和Anthropic的十种模型配置进行了评估,涵盖复杂布局、基金分隔符和多行列名的财务表格。结果显示,无单一模型能达到专业金融应用中无需监督使用的低错误率。表现最佳的Gemini 3.1 Pro在24个不同复杂度的文件上达到82.4%准确率(约每6题错1题),其次为GPT-5.2推理版(80.4%)、Claude Opus 4.6思考模式(80.2%)和Gemini 3 Pro(80.2%)。在最大规模表格(152家公司,8个基金)上,所有模型平均准确率仅为48.6%,远低于最简单文件的86.2%。该难度趋势在所有十种模型中一致,表明这是大模型的共性局限,而非个别缺陷。可靠的财务表格提取可能需要将文档理解与确定性计算分离的架构设计。

原文摘要 · Abstract (English)

While Large Language Models (LLMs) can accelerate text-heavy tasks in alternative investment due diligence, a gap remains in their ability to accurately extract and reason over structured tabular data from complex financial spreadsheets. Progress is held back by the lack of real industry fund portfolio datasets for benchmarking, as private equity data rooms are confidential. To address this, we introduce FinSheet-Bench, a benchmark of synthetic financial portfolio data modeled on real private equity fund structures, designed to evaluate LLM performance on text-serialized spreadsheet question answering and numeric reasoning tasks. Our evaluation of ten model configurations from OpenAI, Google, and Anthropic on financial spreadsheets, including complex layouts, fund dividers, and multi-line column names, reveals that no standalone model achieves error rates low enough for unsupervised use in professional finance applications. The best-performing model, Gemini 3.1 Pro, achieves 82.4% accuracy across twenty-four evaluation files of varying complexity and structural layout (approximately 1 error per 6 questions), followed by GPT-5.2 with reasoning at 80.4%, Claude Opus 4.6 with thinking at 80.2%, and Gemini 3 Pro at 80.2%. Performance degrades substantially on larger, more complex spreadsheets: the largest spreadsheet (152 companies, 8 funds) yields an average accuracy of just 48.6% across all models, compared to 86.2% on the easiest evaluation file. These difficulty patterns are consistent across all ten models, indicating that they reflect LLM limitations rather than idiosyncratic model weaknesses. Reliable financial spreadsheet extraction will likely require architectural approaches that separate document understanding from deterministic computation.

金融分析大模型评测表格推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。