构建中文金融大模型评估体系,量化对比多个模型表现。
FLAME: Financial Large-Language Model Assessment and Metrics Evaluation
- 设计双基准评测集:证书题库与业务场景任务
- 覆盖16000+金融认证题和近百项真实业务任务
- 首次系统评测中文金融大模型,适合研究者参考
大型语言模型(LLM)在自然语言处理领域引发变革,并展现出跨领域的潜力。越来越多面向金融任务的专用大模型被推出,但对其价值进行全面评估仍具挑战。本文提出FLAME——一个面向中文场景的综合性金融大模型评估系统,包含两个核心评测基准:FLAME-Cer与FLAME-Sce。FLAME-Cer涵盖14类权威金融认证,包括CPA、CFA、FRM等,共约16000道经人工审核的精选题目,确保准确性与代表性。FLAME-Sce包含10个主要金融业务场景、21个次要场景及近100项综合应用任务。我们评估了6个代表性模型,包括GPT-4o、GLM-4、ERNIE-4.0、Qwen2.5、XuanYuan3和最新发布的Baichuan4-Finance,结果表明该模型在多数任务中表现最优。通过建立专业、全面的评估体系,FLAME推动了中文金融大模型的发展。参与评测方式见GitHub:https://github.com/FLAME-ruc/FLAME。
原文摘要 · Abstract (English)
LLMs have revolutionized NLP and demonstrated potential across diverse domains. More and more financial LLMs have been introduced for finance-specific tasks, yet comprehensively assessing their value is still challenging. In this paper, we introduce FLAME, a comprehensive financial LLMs evaluation system in Chinese, which includes two core evaluation benchmarks: FLAME-Cer and FLAME-Sce. FLAME-Cer covers 14 types of authoritative financial certifications, including CPA, CFA, and FRM, with a total of approximately 16,000 carefully selected questions. All questions have been manually reviewed to ensure accuracy and representativeness. FLAME-Sce consists of 10 primary core financial business scenarios, 21 secondary financial business scenarios, and a comprehensive evaluation set of nearly 100 tertiary financial application tasks. We evaluate 6 representative LLMs, including GPT-4o, GLM-4, ERNIE-4.0, Qwen2.5, XuanYuan3, and the latest Baichuan4-Finance, revealing Baichuan4-Finance excels other LLMs in most tasks. By establishing a comprehensive and professional evaluation system, FLAME facilitates the advancement of financial LLMs in Chinese contexts. Instructions for participating in the evaluation are available on GitHub: https://github.com/FLAME-ruc/FLAME.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。