为金融行业定制大模型评估体系,解决通用榜单不适用的问题。
Meta-Benchmarks for Financial-Services LLM Evaluation

- 按金融工作活动分类整合452个基准测试,构建跨领域评估框架。
- 用动态权重筛选有效测试,避免过时或饱和的榜单干扰评分。
- 适合金融机构选型、模型治理与合规性评估的决策参考。
公开的大模型排行榜侧重全局平均表现,无法反映金融服务业特有的认知需求:在MMLU-Pro上领先的模型可能在文档驱动的合规推理中表现不佳,擅长编码的模型也可能难以应对多轮客户交互。本文提出一种元基准评估框架,将452个公开报告的基准测试归类至41个O*NET通用工作任务,并聚合为38个涵盖销售、运营、风险与支持工作的BIAN银行业务领域。采用乘法加权机制(区分度×覆盖度×时效性),基于滚动模型窗口计算,自动抑制已饱和的旧测试,保留仍能区分顶尖模型、广泛使用且持续活跃的基准。该权重用于配对式Elo锦标赛中的K因子,生成跨基准可比的工作任务得分,无需原始分数归一化;业务领域得分则为组成任务Elo的加权平均。我们在2026年6月的公开快照中验证该框架,涵盖288个模型及25家机构,详细阐述方法、完整分类体系、设计选择与局限,旨在使该方法可复现,供面临类似选型与治理挑战的机构使用。
原文摘要 · Abstract (English)
Public LLM leaderboards optimise for global average performance and do not capture the specific cognitive demands of financial-services work: a model that leads on MMLU-Pro may underperform on document-grounded compliance reasoning, and a coding leader may handle multi-turn customer interactions poorly. We present a meta-benchmarking framework that organises 452 publicly reported benchmarks into 41 O*NET Generalized Work Activities and aggregates those into 38 BIAN banking business domains spanning sales, operations, risk, and support work. A multiplicative weighting scheme (discrimination x coverage x recency), computed over a rolling model window, rewards benchmarks that still separate the best models, are widely reported, and remain in active use, suppressing saturated legacy tests automatically. These weights scale the K-factor in a pairwise Elo tournament, producing cross-benchmark-comparable work-activity scores without raw score normalisation; business-domain scores are weighted averages of the constituent work-activity Elos. We demonstrate the framework on a point-in-time public snapshot covering 288 models across 25 organisations as of June 2026, and describe the methodology, full taxonomy, design decisions, and limitations with the aim of making the approach reproducible for institutions facing similar selection and governance challenges.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。