arXiv:2505.04110cs.LGcs.CL2025-05被引 1

用真实金融建模竞赛题测试大模型,发现模型在数值推理上表现差。

Alpha Excel Benchmark

  • 将113个金融建模比赛题转为可程序评估的JSON格式
  • 模型在模式识别强但复杂计算能力弱,差距明显
  • 适合关注商业应用的大模型评估者使用

本研究提出一个基于金融建模世界杯(FMWC)Excel竞赛的新基准,用于评估大型语言模型(LLMs)。我们提出一种方法,将113个现有FMWC挑战转化为可程序评估的JSON格式,并利用该数据集对比多个领先LLM的表现。结果表明,不同任务类别间性能差异显著,模型在模式识别任务中表现较好,但在复杂数值推理方面表现不佳。该基准提供了一个标准化框架,用于评估大模型在真实商业场景中的能力,而非抽象学术问题。这项研究推动了人工智能评估的发展,将全球15亿日常使用Microsoft Excel的用户作为有意义的评估指标,弥合了学术基准与实际业务应用之间的差距。

原文摘要 · Abstract (English)

This study presents a novel benchmark for evaluating Large Language Models (LLMs) using challenges derived from the Financial Modeling World Cup (FMWC) Excel competitions. We introduce a methodology for converting 113 existing FMWC challenges into programmatically evaluable JSON formats and use this dataset to compare the performance of several leading LLMs. Our findings demonstrate significant variations in performance across different challenge categories, with models showing specific strengths in pattern recognition tasks but struggling with complex numerical reasoning. The benchmark provides a standardized framework for assessing LLM capabilities in realistic business-oriented tasks rather than abstract academic problems. This research contributes to the growing field of AI benchmarking by establishing proficiency among the 1.5 billion people who daily use Microsoft Excel as a meaningful evaluation metric that bridges the gap between academic AI benchmarks and practical business applications.

大模型评测金融建模Excel应用数值推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。