arXiv:2505.19457cs.AIcs.CE2025-05被引 27

首个面向真实金融场景的中文LLM评估基准,测试模型在财务任务中的实际表现。

BizFinBench: A Business-Driven Real-World Financial Benchmark for Evaluating LLMs

  • 构建覆盖五维度的6781条中文金融任务数据集,含客观与主观评价指标。
  • 发现大模型在计算和推理上领先,但小模型差距显著,信息抽取性能差异最大。
  • 提出新评估方法IteraJudge降低模型自评偏差,适合金融、审计等高精度领域研究者使用。

大型语言模型在通用任务中表现优异,但在金融、法律、医疗等逻辑密集、精度敏感领域评估仍具挑战。为此,我们推出BizFinBench,首个专为真实金融应用设计的LLM评估基准。该基准包含6,781条中文标注查询,涵盖数值计算、推理、信息抽取、预测识别和知识问答五大维度,细分为九个类别,同时提供客观与主观评估指标。我们提出IteraJudge,一种新型的LLM评估方法,可减少模型自评时的偏差。对25个模型(含开源与闭源)的评测显示:无模型在所有任务中占优。具体表现为:数值计算方面,Claude-3.5-Sonnet(63.18)与DeepSeek-R1(64.04)领先,而Qwen2.5-VL-3B仅15.92;推理任务中,闭源模型主导(ChatGPT-o3: 83.58,Gemini-2.0-Flash: 81.15),开源模型落后最高达19.49分;信息抽取性能跨度最大,DeepSeek-R1得71.46,而Qwen3-1.7B仅11.23;预测识别任务表现集中,顶尖模型得分介于39.16至50.00之间。结果表明,当前模型虽能处理常规金融查询,但在需跨概念推理的复杂场景下仍显不足。BizFinBench为未来研究提供了严谨、贴近业务的评估标准。代码与数据集已公开于https://github.com/HiThink-Research/BizFinBench。

原文摘要 · Abstract (English)

Large language models excel in general tasks, yet assessing their reliability in logic-heavy, precision-critical domains like finance, law, and healthcare remains challenging. To address this, we introduce BizFinBench, the first benchmark specifically designed to evaluate LLMs in real-world financial applications. BizFinBench consists of 6,781 well-annotated queries in Chinese, spanning five dimensions: numerical calculation, reasoning, information extraction, prediction recognition, and knowledge-based question answering, grouped into nine fine-grained categories. The benchmark includes both objective and subjective metrics. We also introduce IteraJudge, a novel LLM evaluation method that reduces bias when LLMs serve as evaluators in objective metrics. We benchmark 25 models, including both proprietary and open-source systems. Extensive experiments show that no model dominates across all tasks. Our evaluation reveals distinct capability patterns: (1) In Numerical Calculation, Claude-3.5-Sonnet (63.18) and DeepSeek-R1 (64.04) lead, while smaller models like Qwen2.5-VL-3B (15.92) lag significantly; (2) In Reasoning, proprietary models dominate (ChatGPT-o3: 83.58, Gemini-2.0-Flash: 81.15), with open-source models trailing by up to 19.49 points; (3) In Information Extraction, the performance spread is the largest, with DeepSeek-R1 scoring 71.46, while Qwen3-1.7B scores 11.23; (4) In Prediction Recognition, performance variance is minimal, with top models scoring between 39.16 and 50.00. We find that while current LLMs handle routine finance queries competently, they struggle with complex scenarios requiring cross-concept reasoning. BizFinBench offers a rigorous, business-aligned benchmark for future research. The code and dataset are available at https://github.com/HiThink-Research/BizFinBench.

金融AILLM评估中文数据集基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。