arXiv:2602.17072cs.CL2026-02

构建银行场景数学推理基准,提升大模型算账准确性。

BankMathBench: A Benchmark for Numerical Reasoning in Banking Scenarios

  • 设计三级难度的银行计算任务,覆盖存贷比对与提前还款
  • 微调后模型准确率提升超57个百分点,最高达75%
  • 适合金融AI研发者用于评估和优化模型算术能力

基于大语言模型的聊天机器人正被广泛应用于数字银行,处理存款、储蓄和贷款等产品咨询。然而,这些模型在核心银行业务计算中仍存在高错误率,包括总收益估算、不同利率产品比较以及提前还款条件下的利息计算。这类任务需多步数值推理和对产品上下文的理解,但现有大模型常出现误解产品类型、误用条件或基础运算错误(如指数与等比数列)。现有基准或偏重基础数学,或聚焦金融文档,缺乏对日常银行场景的覆盖。为此,我们提出BankMathBench,一个面向真实银行任务的领域专用数据集,包含三个难度层级:基础(单产品推理)、中级(多产品比较)和高级(多条件场景)。在该数据集上微调开源大模型后,公式生成与数值推理准确率显著提升;采用工具增强微调,平均准确率分别提高57.6%(基础)、75.1%(中级)和62.9%(高级),远超零样本基线。结果表明,BankMathBench是评估和推动大模型在真实银行场景中数值推理能力的有效基准。

原文摘要 · Abstract (English)

Large language models (LLMs)-based chatbots are increasingly being adopted in the financial domain, particularly in digital banking, to handle customer inquiries about products such as deposits, savings, and loans. However, these models still exhibit low accuracy in core banking computations-including total payout estimation, comparison of products with varying interest rates, and interest calculation under early repayment conditions. Such tasks require multi-step numerical reasoning and contextual understanding of banking products, yet existing LLMs often make systematic errors-misinterpreting product types, applying conditions incorrectly, or failing basic calculations involving exponents and geometric progressions. However, such errors have rarely been captured by existing benchmarks. Mathematical datasets focus on fundamental math problems, whereas financial benchmarks primarily target financial documents, leaving everyday banking scenarios underexplored. To address this limitation, we propose BankMathBench, a domain-specific dataset that reflects realistic banking tasks. BankMathBench is organized in three levels of difficulty-basic, intermediate, and advanced-corresponding to single-product reasoning, multi-product comparison, and multi-condition scenarios, respectively. When trained on BankMathBench, open-source LLMs exhibited notable improvements in both formula generation and numerical reasoning accuracy, demonstrating the dataset's effectiveness in enhancing domain-specific reasoning. With tool-augmented fine-tuning, the models achieved average accuracy increases of 57.6%p (basic), 75.1%p (intermediate), and 62.9%p (advanced), representing significant gains over zero-shot baselines. These findings highlight BankMathBench as a reliable benchmark for evaluating and advancing LLMs' numerical reasoning in real-world banking scenarios.

银行智能数学推理大模型评测量化计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。