首个中英文双语金融LLM评测基准,用真实用户数据检验模型实战能力
BizFinBench.v2: Towards Reliable LLMs in Finance via Real-User Data and Offline/Online Bilingual Evaluation
- 基于中美股市真实用户问答构建双语评测集
- GPT-5准确率仅61.5%,远低于84.8%业务要求
- 揭示现有模型在真实金融场景中的关键缺陷
大语言模型在金融应用中日益重要,但现有基准多依赖模拟或通用数据,导致报告性能与实际效果存在显著差距。为此,我们提出BizFinBench.v2,首个基于中、美股票市场真实用户查询-回复数据的离线与在线一体化评测基准。该基准包含8个离线任务和2个在线任务,共28,860个问题。实验显示,GPT-5准确率仅为61.5%,仍未能达到84.8%的实际业务要求。在评估的商用模型中,DeepSeek-R1展现出更优的投资效能。基于真实金融实践的错误分析揭示了现有模型的持续局限性。通过突破以往基准的限制,BizFinBench.v2为推进大语言模型在金融领域的落地提供了可靠依据。数据与代码已公开于https://github.com/HiThink-Research/BizFinBench.v2。
原文摘要 · Abstract (English)
Large language models are becoming increasingly significant in financial applications. Nevertheless, prevailing benchmarks are largely dependent on simulated or generic data, which leads to a significant gap between reported performance and actual efficacy in real-world scenarios. To tackle this challenge, we present BizFinBench.v2, the first integrated offline and online benchmark built upon authentic user query-response data from both Chinese and U.S. equity markets. It comprises 28,860 questions across eight offline and two online tasks. Experimental results show that GPT-5 achieves a mere 61.5% accuracy, still failing to meet the practical business requirement (84.8%). Among the evaluated commercial models, DeepSeek-R1 exhibits superior investment efficacy. Error analysis grounded in real financial practice reveals persistent limitations in existing models. By overcoming the constraints of prior benchmarks, BizFinBench.v2 provides a substantiated foundation for advancing LLM deployment in the financial sector. Our data and code are available at https://github.com/HiThink-Research/BizFinBench.v2.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。